Stronger Normalization-Free Transformers

Published
Source
arXiv
Paper number
099
Field
Architecture
arXiv ID
2512.10938

Key points

  • Traditional normalization layers such as Batch Norm and Layer Norm stabilize deep neural network training but incur overhead from memory access and synchronization and remain sensitive to batch size.
  • Existing normalization-free methods like Dynamic Tanh (DyT) achieve performance comparable to normalization, but they do not systematically surpass it, leaving room for further improvement.
  • There has been a lack of systematic understanding of how the intrinsic properties of pointwise functions affect training dynamics and model performance in normalization-free Transformers.
  • The paper systematically analyzes four intrinsic properties of pointwise functions, zero-centeredness, boundedness, central sensitivity, and monotonicity, to derive design principles for effective normalization-free Transformer components.
  • Based on these principles, the authors extensively explore pointwise functions empirically and identify the error function (erf) as the best-performing design.
  • They develop Dynamic erf (Derf), a pointwise function parameterized as "γ × erf(αx+s)+β" by applying scale, shift, and bias parameters to the error function, which serves as a statistically independent alternative that directly replaces all normalization layers in Transformer architectures.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)