Do Robotic World Models Really Follow Actions? Diagnosing and Aligning Action-Conditioned Generation for Policy Learning

Published
Source
arXiv
Paper number
1011
Field
Robotics
arXiv ID
2608.24885

Key points

  • In addition to demonstrated actions, WorldEcho probes world models with four kinds of non-expert action queries: replay from other states, local perturbations, policy rollouts, and sampling from the feasible space.
  • WorldSync broadens the distribution of action outcomes, anchors intermediate video representations to robot motion with an Action-Forcing Expert, and aligns changes caused by altering actions with changes in the actual future.
  • Across 50 RoboTwin tasks, WorldSync achieved the lowest error accounting for visual integrity, 0.0661, and the highest video pass rate, 84.51%.
  • Using this model as a learned simulator for two rounds of policy improvement raised simulated success rates from 51–52% to 65% and real-robot cup stacking from 48% to 68%, exceeding CtrlWorld by 8–9 and 12 points, respectively.
  • However, its advantage was uneven across individual metrics: Cosmos-Predict2.5 achieved a better raw NDTW of 0.013, and the authors stated that extending the approach to long-duration interactions and multiple robot embodiments remains future work.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)