DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
Ranking
Overall
84
Content
90
Popularity
69
Observed public metrics from 1 member.
Merged summary
TL;DR - DreamX-Phi 1.0 is an action-conditioned video world model that predicts robotic manipulation outcomes from an image, language instruction, and action sequence. It targets faithful arm motion, scene geometry, and object consistency while enabling faster deployment through distillation.
- Encodes per-arm SE(3) transformations in attention using PRoPE-style geometric encoding.
- Adds depth prediction for scene geometry and uses SAM3 masks with a frozen V-JEPA teacher to preserve manipulated objects.
- Distills a multi-step generator into a few-step student for efficient inference.
- Reports first and second place on Tracks 1 and 2, respectively, of the WorldArena 2.0 Challenge.
Sources (1)
DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
Public signals
Hugging Face upvotes 99
TL;DR - DreamX-Phi 1.0 is an action-conditioned video world model that predicts robotic manipulation outcomes from an image, language instruction, and action sequence. It targets faithful arm motion, scene geometry, and object consistency while enabling faster deployment through distillation.
- Encodes per-arm SE(3) transformations in attention using PRoPE-style geometric encoding.
- Adds depth prediction for scene geometry and uses SAM3 masks with a frozen V-JEPA teacher to preserve manipulated objects.
- Distills a multi-step generator into a few-step student for efficient inference.
- Reports first and second place on Tracks 1 and 2, respectively, of the WorldArena 2.0 Challenge.