DiffusionBench: On Holistic Evaluation of Diffusion Transformers

Published
Source
arXiv
Paper number
487
Field
Computer Vision
arXiv ID
2606.24888

Key points

  • It builds NanoGen, a unified training and evaluation framework that can switch from ImageNet to text-to-image training with only 12 lines of configuration changes.
  • After training 21 latent diffusion models, it confirms almost no correlation between ImageNet rank and text-to-image rank, with Pearson coefficients from -0.377 to -0.580.
  • It provides empirical evidence that better ImageNet FID does not translate into better text-to-image quality.
  • It releases a unified codebase that supports RAE, VAE, pixel-space, and MeanFlow.
  • DiffusionBench is a holistic benchmark that combines ImageNet and text-to-image results and is recommended as the default reporting standard for future DiT research.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)