WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
Ranking
Observed public metrics from 1 member.
Merged summary
TL;DR - WorldExam is a hierarchical benchmark that evaluates controllable video generation models as world models, pushing past visual fidelity and explicit instruction-following to test whether generated worlds react plausibly to scene state. It matters because it exposes that no current model combines broad task coverage with consistent performance on inherent reactivity.
- Four evaluation levels — Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity — with 1,474 cases across eight dedicated tasks, unifying camera-, action-, and language-driven paradigms under one protocol.
- The World Reactivity level specifically probes scene-conditioned reactions and goal-directed behaviors that are not explicitly stated in the input, the gap left by prior benchmarks.
- Across 20 models, a clear capability split emerges: camera-driven models win on camera control but lack dynamic-interaction interfaces; action-driven models control subjects precisely yet leave the world unresponsive; language-driven models handle interaction better but follow complex controls less faithfully.
- Core finding: high visual quality and explicit instruction fulfillment are not proxies for inherent reactivity, so world-model claims need reactivity-specific evaluation.
Sources (1)
WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
TL;DR - WorldExam is a hierarchical benchmark that evaluates controllable video generation models as world models, pushing past visual fidelity and explicit instruction-following to test whether generated worlds react plausibly to scene state. It matters because it exposes that no current model combines broad task coverage with consistent performance on inherent reactivity.
- Four evaluation levels — Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity — with 1,474 cases across eight dedicated tasks, unifying camera-, action-, and language-driven paradigms under one protocol.
- The World Reactivity level specifically probes scene-conditioned reactions and goal-directed behaviors that are not explicitly stated in the input, the gap left by prior benchmarks.
- Across 20 models, a clear capability split emerges: camera-driven models win on camera control but lack dynamic-interaction interfaces; action-driven models control subjects precisely yet leave the world unresponsive; language-driven models handle interaction better but follow complex controls less faithfully.
- Core finding: high visual quality and explicit instruction fulfillment are not proxies for inherent reactivity, so world-model claims need reactivity-specific evaluation.