LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies
- Published
- Source
- arXiv
- Paper number
- 426
- Field
- Robotics
- arXiv ID
- 2606.15768
Key points
- LaWM is a 230M-parameter latent world model that reuses the decoder from the latent action model, using 95 percent fewer world-modeling parameters than pixel-space WAMs.
- It predicts latent visual subgoals in a single forward pass and reaches 187 ms through non-iterative inference, up to 24 times faster than pixel-space WAMs.
- The two-stage training pipeline first trains LaWM on about 3,000 hours of robot data and 1,500 hours of egocentric human video, then trains the policy through latent-action distillation.
- In real-world evaluation, it achieves 93.3 percent on Pick-and-Place, 86.7 percent on Drawer Opening, and 90.0 percent on Towel Folding, for an overall average of 90.0 percent and first place among all baselines.
Paper links
External research summaries. These are not HDATF publications or measured product results.