Qwen-Image-Flash: Beyond Objective Design

Published
Source
arXiv
Paper number
358
Field
Computer Vision
arXiv ID
2606.03746

Key points

  • It trains a 4-NFE student model distilled from Qwen-Image-2.0 using distribution matching distillation, or DMD.
  • A counterintuitive data finding emerges: increasing diversity hurts performance, while consistent single-category data transfers better.
  • Step-wise multi-teacher guidance combines complementary strengths across teacher models and stabilizes training.
  • It achieves the best overall performance at a balanced T2I and Edit ratio, and edit supervision also improves generation quality.
  • Qwen-Image-Flash produces high-quality results with only 4 NFE in scenarios such as poster generation.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)