Faster-WAM: Do World Action Models Need Deep Action Modules?
- Published
- Source
- arXiv
- Paper number
- 794
- Field
- AI / General
- arXiv ID
- 2608.02365
Key points
- The main issue is speed. Even Fast-WAM, a representative low-latency model, takes 211.7 ms for one action chunk, about three times the 71.4 ms of π0.5, a VLA model.
- DoT removes the one-to-one alignment between video layers and action layers required by prior MoT methods, so action-head depth can be set independently of backbone depth and can access representations from all backbone layers.
- On the average of four LIBERO task groups, it reaches 98.5%, 0.9 points above Fast-WAM's 97.6%, and gains 2.6 points on LIBERO-Long. It also outperforms π0.5, X-VLA, and Motus, all of which use robotics pretraining.
- Under the same 24GB consumer GPU setting, latency is 66.5 ms for Faster-WAM, 68.2 ms for π0, 71.4 ms for π0.5, 105.3 ms for X-VLA, and 211.7 ms for Fast-WAM. LingBot-VA and Motus could not be measured because they ran out of memory.
- On out-of-distribution generalization in LIBERO-Plus, it reaches 75.0% and beats Fast-WAM's 51.5% by 23.5 points. Camera perturbation rises from 16.4% to 67.9%, sensor noise from 37.7% to 82.7%, and language perturbation from 68.9% to 92.1%.
- Ablation starts at 49.5% and climbs step by step to 60.3% with simple DoT using only the final layer, 66.8% with KV-Fusion, 71.3% with RoPE alignment, and 75.0% after removing text cross-attention. Layer mixing signals are concentrated in middle layers but spread widely across all layers.
Paper links
External research summaries. These are not HDATF publications or measured product results.