Unveiling the Secret of AdaLN-Zero in Diffusion Transformer
- Published
- Source
- arXiv
- Paper number
- 871
- Field
- Computer Vision
- arXiv ID
- 2608.09438
Key points
- It breaks down the performance gains of adaLN-Zero into three factors: SE structure, zero initialization, and progressive updates.
- Zero initialization is the most important factor, and the paper finds that the conditional modulation weights follow a Gaussian distribution after training.
- It proposes adaLN-Gaussian, which initializes with a Gaussian distribution, to validate the analysis and obtain practical gains.
- It proposes SE-adaLN-Zero, inspired by the SE structure, and achieves additional performance gains.
- It confirms consistent improvements and generalization on four datasets, including ImageNet1K.
Paper links
External research summaries. These are not HDATF publications or measured product results.