🛰️ Daily AI Frontier
‹ back to 2026-08-04

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

arXiv cs.CV World Models Yuxue Yang, Shuyao Shang, Jiahe Wang, Zitong Zhou, Liang Tan, Junhan Zeng, Ruizhi Li, Junyan Li, Yu Liu, Xiao Yang, Yong Li, Jun Zhu, Hongsheng Li, Tieniu Tan, Lue Fan, Zhaoxiang Zhang 2026-08-03
Representative image for WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

TL;DR - WorldExam is a hierarchical benchmark that evaluates controllable video generation models as world models, pushing past visual fidelity and explicit instruction-following to test whether generated worlds react plausibly to scene state. It matters because it exposes that no current model combines broad task coverage with consistent performance on inherent reactivity.

  • Four evaluation levels — Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity — with 1,474 cases across eight dedicated tasks, unifying camera-, action-, and language-driven paradigms under one protocol.
  • The World Reactivity level specifically probes scene-conditioned reactions and goal-directed behaviors that are not explicitly stated in the input, the gap left by prior benchmarks.
  • Across 20 models, a clear capability split emerges: camera-driven models win on camera control but lack dynamic-interaction interfaces; action-driven models control subjects precisely yet leave the world unresponsive; language-driven models handle interaction better but follow complex controls less faithfully.
  • Core finding: high visual quality and explicit instruction fulfillment are not proxies for inherent reactivity, so world-model claims need reactivity-specific evaluation.

view merged work →