Masked Visual Actions for Unified World Modeling
Merged summary
TL;DR - Masked Visual Actions turns video models into unified robotic world models by representing actions as partially revealed pixel-space trajectories. One model can predict scene responses, evaluate candidate plans, and infer robot motions for desired object outcomes.
- Revealed robot motion conditions forward-dynamics predictions of scene behavior.
- Revealed target-object motion enables inverse modeling of compatible robot actions.
- A single checkpoint was fine-tuned on 15 hours of masked real and simulated video.
- Imagined rollouts correlate with real execution and improve model-based planning by ranking candidate futures.
Sources (1)
Masked Visual Actions for Unified World Modeling
TL;DR - Masked Visual Actions turns video models into unified robotic world models by representing actions as partially revealed pixel-space trajectories. One model can predict scene responses, evaluate candidate plans, and infer robot motions for desired object outcomes.
- Revealed robot motion conditions forward-dynamics predictions of scene behavior.
- Revealed target-object motion enables inverse modeling of compatible robot actions.
- A single checkpoint was fine-tuned on 15 hours of masked real and simulated video.
- Imagined rollouts correlate with real execution and improve model-based planning by ranking candidate futures.