FlowWAM: Optical Flow as a Unified Action Representation for World Action Models
- Published
- Source
- arXiv
- Paper number
- 619
- Field
- Robotics
- arXiv ID
- 2607.13017
Key points
- It represented actions as optical flow in the same format as RGB video and jointly diffusion-modeled RGB and flow within a shared video generator.
- In policy mode, it predicts actions from flow; in world-model mode, target flow guides future video. Videos without action labels are also used for pretraining.
- Success rates on RoboTwin were 92.94% for Clean and 92.14% for Random, and it achieved an EWMScore of 63.71 on WorldArena.
- By connecting robot policies and future-scene prediction through a single video-friendly representation, it can use motion information from large-scale general video for control learning.
- The reported figures come from RoboTwin, WorldArena, and a limited set of real-robot tasks, so generality across other robot embodiments and contact-centered tasks requires further validation.
Paper links
External research summaries. These are not HDATF publications or measured product results.