🛰️ Daily AI Frontier
‹ back to 2026-08-16

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

arXiv cs.CV Multimodal & Generative DreamX Team, Rui Chen, Xiangxiang Chu, Geng Li, Jifan Li, Qingfeng Shi, Datao Tang, Jing Tang, Jun Wang, Pengfei Zhang 2026-08-13
Representative image for DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

TL;DR - DreamX-Phi 1.0 is an action-conditioned video world model that predicts robotic manipulation outcomes from an image, language instruction, and action sequence. It targets faithful arm motion, scene geometry, and object consistency while enabling faster deployment through distillation.

  • Encodes per-arm SE(3) transformations in attention using PRoPE-style geometric encoding.
  • Adds depth prediction for scene geometry and uses SAM3 masks with a frozen V-JEPA teacher to preserve manipulated objects.
  • Distills a multi-step generator into a few-step student for efficient inference.
  • Reports first and second place on Tracks 1 and 2, respectively, of the WorldArena 2.0 Challenge.

view merged work →