DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
- Published
- Source
- arXiv
- Paper number
- 894
- Field
- Computer Vision
- arXiv ID
- 2608.13489
Key points
- We inject the 3D motion path of a robot arm geometrically into attention, preventing the model from moving the wrong arm even when the video looks plausible.
- A depth prediction branch and SAM3 masks, a technique that separates object regions, reduce mistakes where grasped objects disappear or get swapped during the video.
- We enable fast deployment by distilling the multi-stage generator into a few stages.
- On WorldArena 2.0, Track 1, it ranks first among 31 teams with EWMScore-P 60.65, and it ranks joint second on Track 2 with 67.19%.
- On the WorldArena 1.0 offline evaluation, it scores 76.88, which is 3.24 points above the previous best of 73.64.
Paper links
External research summaries. These are not HDATF publications or measured product results.