SeFi-Image: A Text-to-Image Foundation Model with Semantic-First Diffusion
- Published
- Source
- arXiv
- Paper number
- 477
- Field
- Computer Vision
- arXiv ID
- 2606.22568
Key points
- Semantic-First Diffusion injects semantic encoder features asynchronously and reaches fast convergence even in high-fidelity VAE spaces.
- Training the 5B model requires only 125K A800 hours of compute, which is 10 to 20 percent of the compute used by Z-Image.
- It matches or exceeds Qwen-Image and Z-Image on many benchmarks, including GenEval, DPG, LongTextBench, OneIG, and CVTG-2K.
- Three scales, 1B, 2B, and 5B, plus a DMD2-distilled turbo variant cover different hardware and latency requirements.
- It is trained on 450M image-text pairs plus 28M synthetic text-rendering samples, and the code and weights are released.
Paper links
External research summaries. These are not HDATF publications or measured product results.