DreamWAM: Beyond RGB Future Prediction for World Action Models

Published
Source
arXiv
Paper number
822
Field
Robotics
arXiv ID
2608.04996

Key points

  • It extends future prediction beyond RGB to motion, depth, and semantic information, which substantially improves the robustness of robot action models.
  • The extra information is used only during training, while inference uses RGB alone, so deployment cost does not increase.
  • On standard LIBERO, performance improves slightly from 97.3 percent to 98.4 percent, while on unseen shifts in LIBERO-Plus it improves substantially from 51.4 percent to 63.4 percent.
  • Real-robot experiments show the same pattern, with standard settings improving from 90.8 percent to 96.7 percent and perturbed settings improving from 55.6 percent to 74.4 percent.
  • Among the three auxiliary signals, motion matters most, and depth and semantics work best when injected as residual information.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)