Why Muon Outperforms Adam: A Curvature Perspective
- Published
- Source
- arXiv
- Paper number
- 353
- Field
- Machine Learning
- arXiv ID
- 2606.04662
Key points
- A second-order Taylor expansion decomposes the loss reduction of Muon and Adam. The first-order gains are similar, but the difference appears in the second-order curvature penalty.
- The update norms are similar, but Muon has lower NDS (normalized direction sharpness), so its curvature cost is smaller.
- Controlled experiments with Zipf-PCFG show that data imbalance amplifies Muon's NDS advantage.
- In the middle and later stages of training, Muon's lower NDS comes mainly from reduced within-layer curvature.
- On heterogeneous-curvature second-order problems, Muon allocates update energy evenly across high- and low-curvature directions.
- It is rigorously proven that when curvature heterogeneity is sufficiently large, Muon achieves a lower loss than GD.
Paper links
External research summaries. These are not HDATF publications or measured product results.