🛰️ Daily AI Frontier
‹ back to 2026-08-10

清华大学李升波团队:将JEPA与受控世界模型结合,揭示物理状态与动作转移的可辨识条件

Research World Models

Ranking

Overall 62
Content 75
Popularity 33

Observed public metrics from 1 member.

Representative image for 清华大学李升波团队:将JEPA与受控世界模型结合,揭示物理状态与动作转移的可辨识条件

Merged summary

TL;DR - Tsinghua's iDLab (Shengbo Eben Li) with DiDi's Voyager Lab extends LeCun's JEPA identifiability theory from autonomous to controlled world models, proving when a latent encoder can recover both true physical states and the action-driven transition dynamics (arXiv:2607.22430).

  • Introduces two policy-dependent metrics: representation margin (spectral gap between weakest first-order latent predictive signal and strongest higher-order nonlinear surrogate) and transition margin (weakest residual conditional action variation given state). Identifiability requires both to be strictly positive.
  • Theorem 1: under joint-Gaussian behavior data, stationary linear-Gaussian controlled transitions, invertible observation map, and a standard-Gaussian representation constraint, JEPA training recovers true state and transition up to an orthogonal matrix Q; proof expands the encoder into Hermite components so a positive spectral gap forces only first-order terms to survive.
  • Theorem 2 gives finite-error bounds splitting error into encoder vs. predictor sources; Theorem 3 shows counterfactual prediction error can be amplified to ~ε/transition-margin — a constructible worst case, not a universal upper bound.
  • Experiments on four nonlinear observation maps (spiral, parabolic, sinusoidal, wave) confirm errors fall as margins grow; with zero action coverage, in-distribution error stays low while counterfactual error and A/B matrix estimation degrade, hurting goal-conditioned planning. Practical takeaway: don't judge world models by one-step in-distribution error alone; add exploration noise for offline data collection.

Sources (1)

清华大学李升波团队:将JEPA与受控世界模型结合,揭示物理状态与动作转移的可辨识条件

WeChat: CVer 2026-08-10 arXiv:2607.22430
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-28 14:24:41.102534 UTC

TL;DR - Tsinghua's iDLab (Shengbo Eben Li) with DiDi's Voyager Lab extends LeCun's JEPA identifiability theory from autonomous to controlled world models, proving when a latent encoder can recover both true physical states and the action-driven transition dynamics (arXiv:2607.22430).

  • Introduces two policy-dependent metrics: representation margin (spectral gap between weakest first-order latent predictive signal and strongest higher-order nonlinear surrogate) and transition margin (weakest residual conditional action variation given state). Identifiability requires both to be strictly positive.
  • Theorem 1: under joint-Gaussian behavior data, stationary linear-Gaussian controlled transitions, invertible observation map, and a standard-Gaussian representation constraint, JEPA training recovers true state and transition up to an orthogonal matrix Q; proof expands the encoder into Hermite components so a positive spectral gap forces only first-order terms to survive.
  • Theorem 2 gives finite-error bounds splitting error into encoder vs. predictor sources; Theorem 3 shows counterfactual prediction error can be amplified to ~ε/transition-margin — a constructible worst case, not a universal upper bound.
  • Experiments on four nonlinear observation maps (spiral, parabolic, sinusoidal, wave) confirm errors fall as margins grow; with zero action coverage, in-distribution error stays low while counterfactual error and A/B matrix estimation degrade, hurting goal-conditioned planning. Practical takeaway: don't judge world models by one-step in-distribution error alone; add exploration noise for offline data collection.
item →