Masked Visual Actions for Unified World Modeling
- Published
- Source
- arXiv
- Paper number
- 681
- Field
- Computer Vision
- arXiv ID
- 2607.19343
Key points
- It proposes an interface called Masked Visual Actions, which represents robot actions as pixel-level masks.
- A single checkpoint supports both forward simulation and inverse action generation.
- It works well across different environments and robot embodiments even when it is fine-tuned on only 15 hours of real robot data.
- The imagined rollouts from the video model are strongly correlated with real execution, with r = 0.982.
- Given a desired object motion, the inverse model synthesizes the corresponding robot action and reaches a 90 percent success rate.
Paper links
External research summaries. These are not HDATF publications or measured product results.