MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models
- Published
- Source
- arXiv
- Paper number
- 408
- Field
- Computer Vision
- arXiv ID
- 2606.13515
Key points
- It unifies mask prompting and prediction by using masks both as input, in the form of a first-frame visual prompt, and as output, in the form of future mask prediction.
- Future mask prediction provides object-centric semantic supervision, improving performance over RGB-only methods.
- A first-frame mask prompt resolves language ambiguity and is overwhelmingly better than text coordinate prompts.
- It uses a Mixture of Transformers (MoT) to jointly predict RGB, masks, and actions with unified flow matching.
- It achieves SOTA results of 98.4% on LIBERO and 92.2% on RoboTwin, and 84.3% on clear-language real-robot tasks and 84.9% on ambiguous tasks.
- Without mask prediction, performance drops to 21.6%, showing that the prediction objective is essential for using visual prompts.
Paper links
External research summaries. These are not HDATF publications or measured product results.