Flash-WAM: Modality-Aware Distillation for World Action Models

Published
Source
arXiv
Paper number
355
Field
Machine Learning
arXiv ID
2606.05254

Key points

  • It identifies the structural reason why prior consistency distillation fails in video-action joint diffusion: gradient signals vanish because of an asymmetric noise schedule.
  • It selects modality-specific consistency functions, using linear-gradient scaling for actions and variance-preserving parameterization for video.
  • It delivers a 23 times speedup, reducing per-chunk latency from 8.1 seconds to 348 milliseconds and enabling real-time control.
  • It preserves an 85.5 percent success rate on RoboTwin 2.0 and 95.7 percent on LIBERO under the 1v/2a setting.
  • It achieves a 60 percent success rate on a real-world Unitree G1 humanoid, which is a substantial improvement over the naive LCM baseline at 24 percent.
  • A structural analysis of the consistency function family theoretically explains the limits of gradient scaling.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)