Masked Visual Actions for Unified World Modeling

Published
Source
arXiv
Paper number
681
Field
Computer Vision
arXiv ID
2607.19343

Key points

  • It proposes an interface called Masked Visual Actions, which represents robot actions as pixel-level masks.
  • A single checkpoint supports both forward simulation and inverse action generation.
  • It works well across different environments and robot embodiments even when it is fine-tuned on only 15 hours of real robot data.
  • The imagined rollouts from the video model are strongly correlated with real execution, with r = 0.982.
  • Given a desired object motion, the inverse model synthesizes the corresponding robot action and reaches a 90 percent success rate.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)