🛰️ Daily AI Frontier
‹ back to 2026-08-16

PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

arXiv cs.CV Multimodal & Generative Kaixin Ding, Xi Chen, Minghong Cai, Zhiyuan Xu, Yiyang Wang, Yuxiang Lu, Junyi Li, Shuyang Chen, Yuan Gao, Xin Tao, Pengfei Wan, Hengshuang Zhao 2026-08-13
Representative image for PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

TL;DR - PlayWorld benchmarks video world models using multimodal agents pursuing long-horizon objectives rather than fixed action sequences. Tests across nine models expose persistent weaknesses in spatial consistency and state evolution.

  • Provides 171 objective-driven interactive scenarios.
  • Evaluates geometry consistency, interaction fidelity, and visible and out-of-sight evolution.
  • Also measures basic video quality and action controllability.
  • Agent players enable fairer comparisons when different models require different action sequences to reach the same goal.

view merged work →