Masked Visual Actions for Unified World Modeling
Ranking
Overall
88
Content
95
Popularity
72
Observed public metrics from 1 member.
Merged summary
TL;DR - Masked Visual Actions turns video models into unified robotic world models by representing actions as partially revealed pixel-space trajectories. One model can predict scene responses, evaluate candidate plans, and infer robot motions for desired object outcomes.
- Revealed robot motion conditions forward-dynamics predictions of scene behavior.
- Revealed target-object motion enables inverse modeling of compatible robot actions.
- A single checkpoint was fine-tuned on 15 hours of masked real and simulated video.
- Imagined rollouts correlate with real execution and improve model-based planning by ranking candidate futures.
Sources (1)
Masked Visual Actions for Unified World Modeling
Public signals
Hugging Face upvotes 9 · Semantic Scholar citations 1 · Semantic Scholar influential citations 0
TL;DR - Masked Visual Actions turns video models into unified robotic world models by representing actions as partially revealed pixel-space trajectories. One model can predict scene responses, evaluate candidate plans, and infer robot motions for desired object outcomes.
- Revealed robot motion conditions forward-dynamics predictions of scene behavior.
- Revealed target-object motion enables inverse modeling of compatible robot actions.
- A single checkpoint was fine-tuned on 15 hours of masked real and simulated video.
- Imagined rollouts correlate with real execution and improve model-based planning by ranking candidate futures.