Swift-Image: Exploring the Performance Frontier of Compact Unified Image Generation Models
- Published
- Source
- arXiv
- Paper number
- 962
- Field
- Computer Vision
- arXiv ID
- 2608.20334
Key points
- It built a unified 6B model that handles text-to-image generation, single-image editing, and multi-image editing with one set of weights.
- When the prompt enhancer (PE) converted requests into visual task instructions, the knowledge-intensive benchmark score jumped from 2.49→4.63. The bottleneck in complex generation was intent interpretation rather than rendering.
- Structural pruning retained nearly all performance while reducing the model to 3B, and the few-step distilled model (Turbo) actually achieved a higher overall editing score (4.20 vs 4.16).
- Training required only approximately 243K GPU hours, showing that strong performance is possible without tens of billions of parameters.
- It summarized practical lessons such as 'evolve the training-data distribution to match model capabilities' and 'train conflicting objectives separately before integrating them.'
Paper links
External research summaries. These are not HDATF publications or measured product results.