Latent Action as Intention Enables Efficient Future Imagination for World Action Models
- Published
- Source
- arXiv
- Paper number
- 1008
- Field
- Robotics
- arXiv ID
- 2608.24882
Key points
- A discrete latent-action tokenizer, trained to emphasize regions around hands and manipulators using masks automatically generated by SAM 2, compresses changes in video to form future intentions.
- During training, video, latent actions, and an action expert are learned together. At inference, the future-video branch is removed, and only latent intentions and action chunks are jointly denoised.
- Average success on RoboCasa was 65.6% with limited data and 80.8% with the full dataset, exceeding Fast-WAM under the same conditions by 9.6 and 4.5 points. On LIBERO-Plus, it achieved 74.4% without additional training, 4.0 points above Joint-WAM.
- At 338.5ms per action chunk, it was 42.9% faster than Joint-WAM's 593.1ms, making it favorable for real robot control, where latency can directly lead to failure.
- However, in speed alone it remained slower than Fast-WAM's 196.5ms, which discards future imagination entirely. Without pretraining on egocentric videos lacking action information, the latent-action approach was actually weaker than Joint-WAM under the same conditions.
Paper links
External research summaries. These are not HDATF publications or measured product results.