i1: A Simple and Fully Open Recipe for Strong Text-to-Image Models
- Published
- Source
- arXiv
- Paper number
- 401
- Field
- Computer Vision
- arXiv ID
- 2606.11289
Key points
- A single strong text encoder, T5Gemma-2B, and a large adapter with two transformer blocks are more effective than multiple encoders.
- A dual-stream DiT backbone, long skip connections, and QK-Norm form a strong backbone design, while AdaLN offers little benefit in T2I.
- Training on long captions outperforms short captions, but prompt rewriting is needed at inference time.
- Equal weighting, meaning the same number of images from each dataset, is a strong default for mixing curated datasets.
- i1-3B uses only public data and is on average 29.5 percentage points better than the previous best fully open model across five benchmarks.
- Larger models such as 17B HiDream-I1 and 12B FLUX.1[Dev] are also surpassed, confirming the value of design exploration.
Paper links
External research summaries. These are not HDATF publications or measured product results.