Long-WAM: Scaling the Context of World-Action Models

Published
Source
arXiv
Paper number
1172
Field
Robotics
arXiv ID
2610.10528

Key points

  • Extending the visual history from 0 to 19.2 seconds raised manipulation success from 63.3% to 78.7%, but only for models pretrained autoregressively.
  • They used a two-stage approach: first pretrain a foundation model that predicts the future from the past on 10,000 hours of robot and egocentric video, then adapt it to robot actions.
  • Thanks to streaming encoding and low-precision quantization, an action chunk is generated in 107.4 ms even on a consumer GPU, enabling real-time control.
  • On the task of grasping and stacking cups on a moving conveyor, Long-WAM succeeded 95% of the time, while both comparison baselines failed all 20 trials.
  • Combined with an upstream planning model (GPT family), overall success rose from 31.4% to 54.4%, showing that good planning pays off only with a strong executor.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)