XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?
TL;DR - XEWorld is a controlled cross-embodiment benchmark that tests whether action-conditioned world models for robotic manipulation can render robots they never trained on, and finds they largely behave as 2D visual pattern matchers rather than learned physics simulators.
- The testbed isolates embodiment as the variable by evaluating held-out robots inside physically identical scenes, exposing memorization that training-robot-only evaluation hides.
- Generalization tracks visual similarity, not kinematic similarity; models struggle to map abstract numeric joint actions to coherent visual trajectories and to predict dynamic change from a static initial observation.
- Zero-shot rendering of an unseen embodiment only succeeds with heavily grounded cues — pixel-space actions plus explicit spatial-temporal alignment.
- Few-shot adaptation bypasses the zero-shot barrier but the forced appearance recovery causes catastrophic forgetting of seen embodiments, arguing for architectures that decouple appearance from physical dynamics.