LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies

Published
Source
arXiv
Paper number
426
Field
Robotics
arXiv ID
2606.15768

Key points

  • LaWM is a 230M-parameter latent world model that reuses the decoder from the latent action model, using 95 percent fewer world-modeling parameters than pixel-space WAMs.
  • It predicts latent visual subgoals in a single forward pass and reaches 187 ms through non-iterative inference, up to 24 times faster than pixel-space WAMs.
  • The two-stage training pipeline first trains LaWM on about 3,000 hours of robot data and 1,500 hours of egocentric human video, then trains the policy through latent-action distillation.
  • In real-world evaluation, it achieves 93.3 percent on Pick-and-Place, 86.7 percent on Drawer Opening, and 90.0 percent on Towel Folding, for an overall average of 90.0 percent and first place among all baselines.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)