DreamWAM: Beyond RGB Future Prediction for World Action Models
- Published
- Source
- arXiv
- Paper number
- 822
- Field
- Robotics
- arXiv ID
- 2608.04996
Key points
- It extends future prediction beyond RGB to motion, depth, and semantic information, which substantially improves the robustness of robot action models.
- The extra information is used only during training, while inference uses RGB alone, so deployment cost does not increase.
- On standard LIBERO, performance improves slightly from 97.3 percent to 98.4 percent, while on unseen shifts in LIBERO-Plus it improves substantially from 51.4 percent to 63.4 percent.
- Real-robot experiments show the same pattern, with standard settings improving from 90.8 percent to 96.7 percent and perturbed settings improving from 55.6 percent to 74.4 percent.
- Among the three auxiliary signals, motion matters most, and depth and semantics work best when injected as residual information.
Paper links
External research summaries. These are not HDATF publications or measured product results.