WALL-WM: Carving World Action Modeling at the Event Joints

Published
Source
arXiv
Paper number
351
Field
Robotics
arXiv ID
2606.01955

Key points

  • It shifts the paradigm from chunk-centric optimization to event-grounded optimization, aligning natural boundaries across language, video, and action.
  • It builds a data ecosystem made up of event captions and cluster-balanced sampling.
  • It uses dual inference modes: Event mode for variable-length event execution and Unified mode for fixed chunks with Staircase Decoding.
  • It provides a scalable training recipe built on large-scale pretraining infrastructure with the Muon optimizer.
  • It demonstrates strong generalization across language, scene, and task settings.
  • It outperforms the existing WAM on both physical video prediction and executable control.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)