DC-WAM: Dynamic-Centric Visual Supervision and Reasoning for World-Action Models

Published
Source
arXiv
Paper number
756
Field
Robotics
arXiv ID
2607.25918

Key points

  • We redistribute visual supervision from appearance reconstruction to interaction-centered dynamics in RGB video prediction, so PSNR drops but policy performance improves.
  • A temporal difference flow-matching objective and tracker-guided weighting focus on robot-object contact motion rather than static background and texture.
  • The DynaRoute attention bias predicts dynamic-relevant tokens and steers visual attention toward them.
  • On LIBERO-Plus OOD, it reaches 60.9%, a 9.4-point gain over FastWAM, and reduces the ID-OOD gap from 46.1 to 37.2.
  • On real robot bimanual tasks, it achieves up to a 30-point improvement under lighting changes.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)