Long-WAM: Scaling the Context of World-Action Models
- Published
- Source
- arXiv
- Paper number
- 1172
- Field
- Robotics
- arXiv ID
- 2610.10528
Key points
- Extending the visual history from 0 to 19.2 seconds raised manipulation success from 63.3% to 78.7%, but only for models pretrained autoregressively.
- They used a two-stage approach: first pretrain a foundation model that predicts the future from the past on 10,000 hours of robot and egocentric video, then adapt it to robot actions.
- Thanks to streaming encoding and low-precision quantization, an action chunk is generated in 107.4 ms even on a consumer GPU, enabling real-time control.
- On the task of grasping and stacking cups on a moving conveyor, Long-WAM succeeded 95% of the time, while both comparison baselines failed all 20 trials.
- Combined with an upstream planning model (GPT family), overall success rose from 31.4% to 54.4%, showing that good planning pays off only with a strong executor.
Paper links
External research summaries. These are not HDATF publications or measured product results.