i1: A Simple and Fully Open Recipe for Strong Text-to-Image Models

Published
Source
arXiv
Paper number
401
Field
Computer Vision
arXiv ID
2606.11289

Key points

  • A single strong text encoder, T5Gemma-2B, and a large adapter with two transformer blocks are more effective than multiple encoders.
  • A dual-stream DiT backbone, long skip connections, and QK-Norm form a strong backbone design, while AdaLN offers little benefit in T2I.
  • Training on long captions outperforms short captions, but prompt rewriting is needed at inference time.
  • Equal weighting, meaning the same number of images from each dataset, is a strong default for mixing curated datasets.
  • i1-3B uses only public data and is on average 29.5 percentage points better than the previous best fully open model across five benchmarks.
  • Larger models such as 17B HiDream-I1 and 12B FLUX.1[Dev] are also surpassed, confirming the value of design exploration.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)