Abra: Scaling Diffusion Image Training

Published
Source
arXiv
Paper number
945
Field
Machine Learning
arXiv ID
2608.17286

Key points

  • A controlled model family spanning 10^19 to 10^22 FLOPs was used to measure the compute-optimal point for text-to-image diffusion models.
  • The optimum is about 200 image tokens per parameter, ten times the Chinchilla rule for language models at 20 tokens per parameter.
  • Diffusion models are robust to overtraining, so when compute remains available, using more data is preferable to increasing model size.
  • Generative-quality metrics such as FID and CLIPScore, representation quality, and optimal CFG settings also vary predictably with compute.
  • The compute-optimal token count rises with resolution, making high-resolution training compute-bound rather than data-bound.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)