Swift-Image: Exploring the Performance Frontier of Compact Unified Image Generation Models

Published
Source
arXiv
Paper number
962
Field
Computer Vision
arXiv ID
2608.20334

Key points

  • It built a unified 6B model that handles text-to-image generation, single-image editing, and multi-image editing with one set of weights.
  • When the prompt enhancer (PE) converted requests into visual task instructions, the knowledge-intensive benchmark score jumped from 2.49→4.63. The bottleneck in complex generation was intent interpretation rather than rendering.
  • Structural pruning retained nearly all performance while reducing the model to 3B, and the few-step distilled model (Turbo) actually achieved a higher overall editing score (4.20 vs 4.16).
  • Training required only approximately 243K GPU hours, showing that strong performance is possible without tens of billions of parameters.
  • It summarized practical lessons such as 'evolve the training-data distribution to match model capabilities' and 'train conflicting objectives separately before integrating them.'

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)