Flash-WAM: Modality-Aware Distillation for World Action Models
- Published
- Source
- arXiv
- Paper number
- 355
- Field
- Machine Learning
- arXiv ID
- 2606.05254
Key points
- It identifies the structural reason why prior consistency distillation fails in video-action joint diffusion: gradient signals vanish because of an asymmetric noise schedule.
- It selects modality-specific consistency functions, using linear-gradient scaling for actions and variance-preserving parameterization for video.
- It delivers a 23 times speedup, reducing per-chunk latency from 8.1 seconds to 348 milliseconds and enabling real-time control.
- It preserves an 85.5 percent success rate on RoboTwin 2.0 and 95.7 percent on LIBERO under the 1v/2a setting.
- It achieves a 60 percent success rate on a real-world Unitree G1 humanoid, which is a substantial improvement over the naive LCM baseline at 24 percent.
- A structural analysis of the consistency function family theoretically explains the limits of gradient scaling.
Paper links
External research summaries. These are not HDATF publications or measured product results.