Faster-WAM: Do World Action Models Need Deep Action Modules?

Published
Source
arXiv
Paper number
794
Field
AI / General
arXiv ID
2608.02365

Key points

  • The main issue is speed. Even Fast-WAM, a representative low-latency model, takes 211.7 ms for one action chunk, about three times the 71.4 ms of π0.5, a VLA model.
  • DoT removes the one-to-one alignment between video layers and action layers required by prior MoT methods, so action-head depth can be set independently of backbone depth and can access representations from all backbone layers.
  • On the average of four LIBERO task groups, it reaches 98.5%, 0.9 points above Fast-WAM's 97.6%, and gains 2.6 points on LIBERO-Long. It also outperforms π0.5, X-VLA, and Motus, all of which use robotics pretraining.
  • Under the same 24GB consumer GPU setting, latency is 66.5 ms for Faster-WAM, 68.2 ms for π0, 71.4 ms for π0.5, 105.3 ms for X-VLA, and 211.7 ms for Fast-WAM. LingBot-VA and Motus could not be measured because they ran out of memory.
  • On out-of-distribution generalization in LIBERO-Plus, it reaches 75.0% and beats Fast-WAM's 51.5% by 23.5 points. Camera perturbation rises from 16.4% to 67.9%, sensor noise from 37.7% to 82.7%, and language perturbation from 68.9% to 92.1%.
  • Ablation starts at 49.5% and climbs step by step to 60.3% with simple DoT using only the final layer, 66.8% with KV-Fusion, 71.3% with RoPE alignment, and 75.0% after removing text cross-attention. Layer mixing signals are concentrated in middle layers but spread widely across all layers.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)