WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory

Published
Source
arXiv
Paper number
685
Field
Robotics
arXiv ID
2607.18840

Key points

  • It proposes a reasoning-augmented memory structure that combines short-term visual memory and long-term event memory.
  • It builds a dataset of 5 million event segments called ManipEvent-5M.
  • It can control robots through multiple modalities, including text, goal images, and demonstration videos.
  • On a real dual-arm robot, it achieves 75 percent success on long-horizon tasks and 86.7 percent on fine-grained commands.
  • Its generalization to new object categories improves sharply compared with prior work, from 43.3 percent to 60 percent OOD.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)