Do Robotic World Models Really Follow Actions? Diagnosing and Aligning Action-Conditioned Generation for Policy Learning
TL;DR - WorldEcho reveals that robotic world models often fail to follow valid off-expert actions, limiting their reliability as policy-learning simulators. The proposed WorldSync training framework improves action-conditioned generation and supports more successful iterative policy improvement.
- WorldEcho evaluates action following beyond expert demonstrations using visual integrity and SE(3) trajectory alignment.
- Existing models handle expert actions reasonably but may ignore diverse off-expert commands or generate visually invalid rollouts.
- WorldSync targets distributional coverage, grounding video representations in robot dynamics, and alignment of predicted intervention effects with real futures.
- Experiments on RoboTwin and real robots show improved diagnostic metrics and higher policy success rates.