Modality-Autoregressive World-Action Models

Published
Source
arXiv
Paper number
1092
Field
Robotics
arXiv ID
2609.17524

Key points

  • This paper is the first to propose a 'modality-autoregressive' structure that predicts the future as multiple information types in sequence and generates actions last.
  • It showed that predicting point tracks (motion), DINO features (semantics), and depth (geometry) improves performance, while adding RGB video prediction gives no consistent benefit.
  • It achieved a similar or better success rate (75% vs. 72%) than a 6B pretrained model with roughly 20x less training compute.
  • It beat existing methods on three real-world bimanual robot tasks (cup stacking, towel folding, drawer organization) and improved further when human video data was mixed in.
  • Training from scratch without pretraining enables controlled comparisons, which is a significant lesson for experimental design in this field.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)