🛰️ Daily AI Frontier
‹ back to 2026-08-16

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

Research Multimodal & Generative

Ranking

Overall 84
Content 90
Popularity 69

Observed public metrics from 1 member.

Representative image for DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

Merged summary

TL;DR - DreamX-Phi 1.0 is an action-conditioned video world model that predicts robotic manipulation outcomes from an image, language instruction, and action sequence. It targets faithful arm motion, scene geometry, and object consistency while enabling faster deployment through distillation.

  • Encodes per-arm SE(3) transformations in attention using PRoPE-style geometric encoding.
  • Adds depth prediction for scene geometry and uses SAM3 masks with a frozen V-JEPA teacher to preserve manipulated objects.
  • Distills a multi-step generator into a few-step student for efficient inference.
  • Reports first and second place on Tracks 1 and 2, respectively, of the WorldArena 2.0 Challenge.

Sources (1)

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

arXiv cs.CV DreamX Team, Rui Chen, Xiangxiang Chu, Geng Li, Jifan Li, Qingfeng Shi, Datao Tang, Jing Tang, Jun Wang, Pengfei Zhang 2026-08-13 arXiv:2608.13489
Public signals Hugging Face upvotes 99
Providers: Hugging Face · Upvotes 99 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-15 14:32:55.069763 UTC

TL;DR - DreamX-Phi 1.0 is an action-conditioned video world model that predicts robotic manipulation outcomes from an image, language instruction, and action sequence. It targets faithful arm motion, scene geometry, and object consistency while enabling faster deployment through distillation.

  • Encodes per-arm SE(3) transformations in attention using PRoPE-style geometric encoding.
  • Adds depth prediction for scene geometry and uses SAM3 masks with a frozen V-JEPA teacher to preserve manipulated objects.
  • Distills a multi-step generator into a few-step student for efficient inference.
  • Reports first and second place on Tracks 1 and 2, respectively, of the WorldArena 2.0 Challenge.
item →