On the Principles Behind Neural Network Optimizers

Published
Source
arXiv
Paper number
936
Field
Machine Learning
arXiv ID
2608.16760

Key points

  • Adam is neither always convergent nor always divergent, and the paper gives a safe region that depends on the problem and the number of mini-batches, with smaller batches requiring a larger second-moment coefficient beta2.
  • The Hessian of transformers becomes almost block diagonal during training, and the curvature differs greatly across blocks. The authors explain this as the reason Adam, with coordinate-wise learning rates, can be better than SGD.
  • They analyze simple neural networks and random matrix theory to show that consecutive large matrix multiplications create this Hessian structure, and this view also leads to a design that splits Muon correction across attention heads.
  • Adam-mini replaces more than 99.9% of the second-moment values with block-wise scalars, which cuts optimizer-state memory by 50%. By comparison, Adam needs about 104 GB just for the two moment states in a 13B-parameter model.
  • When training Llama 2 architectures from 39M to 1B parameters on C4, Adam-mini matches AdamW validation loss. However, this is a shortened version with no proofs, and rigorous analysis of attention and Mixture-of-Experts structures is left for future work.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)