Research
DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
arXiv:2608.13489v1 Announce Type: new Abstract: We present extbf{DreamX-Phi 1.0}, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction
arXiv:2608.13489v1 Announce Type: new Abstract: We present extbf{DreamX-Phi 1.0}, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effector poses and gripper states, predicts the resulting future observations. Yet realism alone does not guarantee faithfulness: a convincing rollout can still move the wrong arm or lose the manipulated object. To ensure the prediction respects each arm's commanded path, we inject per-arm SE(3) transformations into attention via extbf{PRoPE-style geometric encoding}, preserving arm identity and rigid-motion structure. Action control alone does not fully constrain scene geometry or the evolution of small manipulated objects. We therefore add a lightweight extbf{depth branch} for scene-level geometry and use extbf{SAM3 masks} with a frozen extbf{V-JEPA teacher} to maintain object consistency throughout grasping. We further distill the multi-step generator into a few-step student via distribution-matching distillation for efficient deployment. At the time of writing, model{} achieves first place on Track1 and second place on Track2 of the WorldArena~2.0 Challenge. Our model and code will be publicly available.
Related
- Wonder: Video World Model Done Better
- ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
- MemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action Models
- StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation
- SPIRAL: Self-Evolving Action-Conditioned Video Generation via Reflective Planning Agents
Source: arXiv cs.CV | 2026-08-14