Unveiling the Secret of AdaLN-Zero in Diffusion Transformer

Published
Source
arXiv
Paper number
871
Field
Computer Vision
arXiv ID
2608.09438

Key points

  • It breaks down the performance gains of adaLN-Zero into three factors: SE structure, zero initialization, and progressive updates.
  • Zero initialization is the most important factor, and the paper finds that the conditional modulation weights follow a Gaussian distribution after training.
  • It proposes adaLN-Gaussian, which initializes with a Gaussian distribution, to validate the analysis and obtain practical gains.
  • It proposes SE-adaLN-Zero, inspired by the SE structure, and achieves additional performance gains.
  • It confirms consistent improvements and generalization on four datasets, including ImageNet1K.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)