On the Principles Behind Neural Network Optimizers
- Published
- Source
- arXiv
- Paper number
- 936
- Field
- Machine Learning
- arXiv ID
- 2608.16760
Key points
- Adam is neither always convergent nor always divergent, and the paper gives a safe region that depends on the problem and the number of mini-batches, with smaller batches requiring a larger second-moment coefficient beta2.
- The Hessian of transformers becomes almost block diagonal during training, and the curvature differs greatly across blocks. The authors explain this as the reason Adam, with coordinate-wise learning rates, can be better than SGD.
- They analyze simple neural networks and random matrix theory to show that consecutive large matrix multiplications create this Hessian structure, and this view also leads to a design that splits Muon correction across attention heads.
- Adam-mini replaces more than 99.9% of the second-moment values with block-wise scalars, which cuts optimizer-state memory by 50%. By comparison, Adam needs about 104 GB just for the two moment states in a 13B-parameter model.
- When training Llama 2 architectures from 39M to 1B parameters on C4, Adam-mini matches AdamW validation loss. However, this is a shortened version with no proofs, and rigorous analysis of attention and Mixture-of-Experts structures is left for future work.
Paper links
External research summaries. These are not HDATF publications or measured product results.