DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

Published
Source
arXiv
Paper number
894
Field
Computer Vision
arXiv ID
2608.13489

Key points

  • We inject the 3D motion path of a robot arm geometrically into attention, preventing the model from moving the wrong arm even when the video looks plausible.
  • A depth prediction branch and SAM3 masks, a technique that separates object regions, reduce mistakes where grasped objects disappear or get swapped during the video.
  • We enable fast deployment by distilling the multi-stage generator into a few stages.
  • On WorldArena 2.0, Track 1, it ranks first among 31 teams with EWMScore-P 60.65, and it ranks joint second on Track 2 with 67.19%.
  • On the WorldArena 1.0 offline evaluation, it scores 76.88, which is 3.24 points above the previous best of 73.64.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)