Latent Action as Intention Enables Efficient Future Imagination for World Action Models

Published
Source
arXiv
Paper number
1008
Field
Robotics
arXiv ID
2608.24882

Key points

  • A discrete latent-action tokenizer, trained to emphasize regions around hands and manipulators using masks automatically generated by SAM 2, compresses changes in video to form future intentions.
  • During training, video, latent actions, and an action expert are learned together. At inference, the future-video branch is removed, and only latent intentions and action chunks are jointly denoised.
  • Average success on RoboCasa was 65.6% with limited data and 80.8% with the full dataset, exceeding Fast-WAM under the same conditions by 9.6 and 4.5 points. On LIBERO-Plus, it achieved 74.4% without additional training, 4.0 points above Joint-WAM.
  • At 338.5ms per action chunk, it was 42.9% faster than Joint-WAM's 593.1ms, making it favorable for real robot control, where latency can directly lead to failure.
  • However, in speed alone it remained slower than Fast-WAM's 196.5ms, which discards future imagination entirely. Without pretraining on egocentric videos lacking action information, the latent-action approach was actually weaker than Joint-WAM under the same conditions.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)