WAM4D: Fast 4D World Action Model via Spatial Register Tokens
- Published
- Source
- arXiv
- Paper number
- 421
- Field
- Computer Vision
- arXiv ID
- 2606.14048
Key points
- The spatial register token handles depth readout only during training and is removed at inference time, so it adds no latency cost.
- Causal Mixture Attention defines modality-specific visibility among video, action, and geometry tokens in the MoT backbone.
- On RoboTwin 2.0, it improves by 7.5 points over pi0 and by 6.1 points over LingBot-VA on the selected 10 tasks.
- Its inference latency is 525 ms, which is slower than the 64 ms of VLA but still fast among WAM systems, and removing the depth branch leaves it at 9.71 GiB of VRAM.
- On real-world long-horizon four-step manipulation tasks, it improves by 13.4 points over pi0.
- The trainable pretrained geometric head is more efficient when it is unidirectional rather than bidirectional, and both beat random initialization.
Paper links
External research summaries. These are not HDATF publications or measured product results.