DiffusionBench: On Holistic Evaluation of Diffusion Transformers
- Published
- Source
- arXiv
- Paper number
- 487
- Field
- Computer Vision
- arXiv ID
- 2606.24888
Key points
- It builds NanoGen, a unified training and evaluation framework that can switch from ImageNet to text-to-image training with only 12 lines of configuration changes.
- After training 21 latent diffusion models, it confirms almost no correlation between ImageNet rank and text-to-image rank, with Pearson coefficients from -0.377 to -0.580.
- It provides empirical evidence that better ImageNet FID does not translate into better text-to-image quality.
- It releases a unified codebase that supports RAE, VAE, pixel-space, and MeanFlow.
- DiffusionBench is a holistic benchmark that combines ImageNet and text-to-image results and is recommended as the default reporting standard for future DiT research.
Paper links
External research summaries. These are not HDATF publications or measured product results.