🛰️ Daily AI Frontier
‹ back to 2026-08-27

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

Research Multimodal & Generative

Ranking

Overall 90
Content 100
Popularity 68

Observed public metrics from 1 member.

Representative image for PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

Merged summary

TL;DR - PAWBench evaluates whether video generators used as world models reproduce the probability distribution of valid physical outcomes, rather than merely generating plausible individual trajectories. Across 50 scenarios and 11 systems, no model consistently recovered both reference probabilities and the full range of valid behaviors.

  • Formalizes “probabilistic alignment” as a distribution-level criterion for stochastic world models.
  • Introduces PAWEval, which converts repeated video rollouts into empirical distributions over physical outcomes.
  • Tests whether language prompts, initial-noise sampling, or model training can reshape predicted outcome distributions.
  • Reveals a significant gap between plausible video generation and distributionally accurate world modeling.

Sources (1)

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

arXiv cs.CV Yuandong Pu, Le Zhuo, Sayak Paul, Gabriel Jorge Menezes, Avram Đorđević, Shiyang Li, Yifan Zhou, Bin Fu, Wenlong Zhang, Junjun He, Yu Qiao, Yihao Liu, Jingbo Xing, Xi Chen 2026-08-27 arXiv:2608.27345
Public signals Hugging Face upvotes 75
Providers: Hugging Face · Upvotes 75 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:27:40.012165 UTC

TL;DR - PAWBench evaluates whether video generators used as world models reproduce the probability distribution of valid physical outcomes, rather than merely generating plausible individual trajectories. Across 50 scenarios and 11 systems, no model consistently recovered both reference probabilities and the full range of valid behaviors.

  • Formalizes “probabilistic alignment” as a distribution-level criterion for stochastic world models.
  • Introduces PAWEval, which converts repeated video rollouts into empirical distributions over physical outcomes.
  • Tests whether language prompts, initial-noise sampling, or model training can reshape predicted outcome distributions.
  • Reveals a significant gap between plausible video generation and distributionally accurate world modeling.
item →