Do Robotic World Models Really Follow Actions? Diagnosing and Aligning Action-Conditioned Generation for Policy Learning
- Published
- Source
- arXiv
- Paper number
- 1011
- Field
- Robotics
- arXiv ID
- 2608.24885
Key points
- In addition to demonstrated actions, WorldEcho probes world models with four kinds of non-expert action queries: replay from other states, local perturbations, policy rollouts, and sampling from the feasible space.
- WorldSync broadens the distribution of action outcomes, anchors intermediate video representations to robot motion with an Action-Forcing Expert, and aligns changes caused by altering actions with changes in the actual future.
- Across 50 RoboTwin tasks, WorldSync achieved the lowest error accounting for visual integrity, 0.0661, and the highest video pass rate, 84.51%.
- Using this model as a learned simulator for two rounds of policy improvement raised simulated success rates from 51–52% to 65% and real-robot cup stacking from 48% to 68%, exceeding CtrlWorld by 8–9 and 12 points, respectively.
- However, its advantage was uneven across individual metrics: Cosmos-Predict2.5 achieved a better raw NDTW of 0.013, and the authors stated that extending the approach to long-duration interactions and multiple robot embodiments remains future work.
Paper links
External research summaries. These are not HDATF publications or measured product results.