Stronger Normalization-Free Transformers
- Published
- Source
- arXiv
- Paper number
- 099
- Field
- Architecture
- arXiv ID
- 2512.10938
Key points
- Traditional normalization layers such as Batch Norm and Layer Norm stabilize deep neural network training but incur overhead from memory access and synchronization and remain sensitive to batch size.
- Existing normalization-free methods like Dynamic Tanh (DyT) achieve performance comparable to normalization, but they do not systematically surpass it, leaving room for further improvement.
- There has been a lack of systematic understanding of how the intrinsic properties of pointwise functions affect training dynamics and model performance in normalization-free Transformers.
- The paper systematically analyzes four intrinsic properties of pointwise functions, zero-centeredness, boundedness, central sensitivity, and monotonicity, to derive design principles for effective normalization-free Transformer components.
- Based on these principles, the authors extensively explore pointwise functions empirically and identify the error function (erf) as the best-performing design.
- They develop Dynamic erf (Derf), a pointwise function parameterized as "γ × erf(αx+s)+β" by applying scale, shift, and bias parameters to the error function, which serves as a statistically independent alternative that directly replaces all normalization layers in Transformer architectures.
Paper links
External research summaries. These are not HDATF publications or measured product results.