MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models

Published
Source
arXiv
Paper number
408
Field
Computer Vision
arXiv ID
2606.13515

Key points

  • It unifies mask prompting and prediction by using masks both as input, in the form of a first-frame visual prompt, and as output, in the form of future mask prediction.
  • Future mask prediction provides object-centric semantic supervision, improving performance over RGB-only methods.
  • A first-frame mask prompt resolves language ambiguity and is overwhelmingly better than text coordinate prompts.
  • It uses a Mixture of Transformers (MoT) to jointly predict RGB, masks, and actions with unified flow matching.
  • It achieves SOTA results of 98.4% on LIBERO and 92.2% on RoboTwin, and 84.3% on clear-language real-robot tasks and 84.9% on ambiguous tasks.
  • Without mask prediction, performance drops to 21.6%, showing that the prediction objective is essential for using visual prompts.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)