Modality-Autoregressive World-Action Models
- Published
- Source
- arXiv
- Paper number
- 1092
- Field
- Robotics
- arXiv ID
- 2609.17524
Key points
- This paper is the first to propose a 'modality-autoregressive' structure that predicts the future as multiple information types in sequence and generates actions last.
- It showed that predicting point tracks (motion), DINO features (semantics), and depth (geometry) improves performance, while adding RGB video prediction gives no consistent benefit.
- It achieved a similar or better success rate (75% vs. 72%) than a 6B pretrained model with roughly 20x less training compute.
- It beat existing methods on three real-world bimanual robot tasks (cup stacking, towel folding, drawer organization) and improved further when human video data was mixed in.
- Training from scratch without pretraining enables controlled comparisons, which is a significant lesson for experimental design in this field.
Paper links
External research summaries. These are not HDATF publications or measured product results.