WAM4D: Fast 4D World Action Model via Spatial Register Tokens

Published
Source
arXiv
Paper number
421
Field
Computer Vision
arXiv ID
2606.14048

Key points

  • The spatial register token handles depth readout only during training and is removed at inference time, so it adds no latency cost.
  • Causal Mixture Attention defines modality-specific visibility among video, action, and geometry tokens in the MoT backbone.
  • On RoboTwin 2.0, it improves by 7.5 points over pi0 and by 6.1 points over LingBot-VA on the selected 10 tasks.
  • Its inference latency is 525 ms, which is slower than the 64 ms of VLA but still fast among WAM systems, and removing the depth branch leaves it at 9.71 GiB of VRAM.
  • On real-world long-horizon four-step manipulation tasks, it improves by 13.4 points over pi0.
  • The trainable pretrained geometric head is more efficient when it is unidirectional rather than bidirectional, and both beat random initialization.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)