🛰️ Daily AI Frontier
‹ back to 2026-07-22

Masked Visual Actions for Unified World Modeling

arXiv cs.CV Multimodal & Generative Hadi Alzayer, Wenlong Huang, Haonan Chen, Christopher Luey, Lvmin Zhang, Maneesh Agrawala, Gordon Wetzstein, Li Fei-Fei, Yilun Du, Jiajun Wu, Jia-Bin Huang 2026-07-21

TL;DR - Masked Visual Actions turns video models into unified robotic world models by representing actions as partially revealed pixel-space trajectories. One model can predict scene responses, evaluate candidate plans, and infer robot motions for desired object outcomes.

  • Revealed robot motion conditions forward-dynamics predictions of scene behavior.
  • Revealed target-object motion enables inverse modeling of compatible robot actions.
  • A single checkpoint was fine-tuned on 15 hours of masked real and simulated video.
  • Imagined rollouts correlate with real execution and improve model-based planning by ranking candidate futures.

view merged work →