Parallel Decoding Distillation for Fast Image and Video Generation

Published
Source
arXiv
Paper number
753
Field
Computer Vision
arXiv ID
2607.26004

Key points

  • It trained a 'parallel decoder' that predicts multiple denoising steps in parallel from a single forward pass, speeding up generation.
  • It solved the mode collapse problem by using only a simple regression objective, without VSD or GAN losses.
  • On Wan2.1 14B text-to-video, it ranked first on VBench and outperformed the baseline on diversity metrics.
  • On Qwen-Image 20B, it achieved SOTA scores with only 4 to 8 NFE.
  • On the LTX-2.3 22B model, it can generate a 720p 10-second video with audio in only 8 NFE.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)