Parallel Decoding Distillation for Fast Image and Video Generation
- Published
- Source
- arXiv
- Paper number
- 753
- Field
- Computer Vision
- arXiv ID
- 2607.26004
Key points
- It trained a 'parallel decoder' that predicts multiple denoising steps in parallel from a single forward pass, speeding up generation.
- It solved the mode collapse problem by using only a simple regression objective, without VSD or GAN losses.
- On Wan2.1 14B text-to-video, it ranked first on VBench and outperformed the baseline on diversity metrics.
- On Qwen-Image 20B, it achieved SOTA scores with only 4 to 8 NFE.
- On the LTX-2.3 22B model, it can generate a 720p 10-second video with audio in only 8 NFE.
Paper links
External research summaries. These are not HDATF publications or measured product results.