FlowWAM: Optical Flow as a Unified Action Representation for World Action Models

Published
Source
arXiv
Paper number
619
Field
Robotics
arXiv ID
2607.13017

Key points

  • It represented actions as optical flow in the same format as RGB video and jointly diffusion-modeled RGB and flow within a shared video generator.
  • In policy mode, it predicts actions from flow; in world-model mode, target flow guides future video. Videos without action labels are also used for pretraining.
  • Success rates on RoboTwin were 92.94% for Clean and 92.14% for Random, and it achieved an EWMScore of 63.71 on WorldArena.
  • By connecting robot policies and future-scene prediction through a single video-friendly representation, it can use motion information from large-scale general video for control learning.
  • The reported figures come from RoboTwin, WorldArena, and a limited set of real-robot tasks, so generality across other robot embodiments and contact-centered tasks requires further validation.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)