SeFi-Image: A Text-to-Image Foundation Model with Semantic-First Diffusion

Published
Source
arXiv
Paper number
477
Field
Computer Vision
arXiv ID
2606.22568

Key points

  • Semantic-First Diffusion injects semantic encoder features asynchronously and reaches fast convergence even in high-fidelity VAE spaces.
  • Training the 5B model requires only 125K A800 hours of compute, which is 10 to 20 percent of the compute used by Z-Image.
  • It matches or exceeds Qwen-Image and Z-Image on many benchmarks, including GenEval, DPG, LongTextBench, OneIG, and CVTG-2K.
  • Three scales, 1B, 2B, and 5B, plus a DMD2-distilled turbo variant cover different hardware and latency requirements.
  • It is trained on 450M image-text pairs plus 28M synthetic text-rendering samples, and the code and weights are released.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)