MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation

Published
Source
arXiv
Paper number
388
Field
Robotics
arXiv ID
2606.09215

Key points

  • It uses intermediate denoising features from a video DiT as policy conditioning, which makes real-time execution possible without generating full videos.
  • Instead of splitting the upper and lower body, it controls the whole body in a single latent space using a unified motion latent variable.
  • Its three-stage training pipeline consists of 2,136 hours of egocentric video pretraining, cross-embodiment action post-training, and whole-body fine-tuning.
  • On the Unitree G1, it achieves a 76.1 percent average success rate across 9 tasks, which is more than 32 percentage points above the strongest VLA baseline at 43.9 percent.
  • It enables task-driven foot interaction, such as kicking a ball or pressing a pedal, which split policies cannot perform.
  • It runs in real time at 4.9 Hz on an NVIDIA A100, about seven times faster than Cosmos Policy at 0.7 Hz.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)