Why Muon Outperforms Adam: A Curvature Perspective

Published
Source
arXiv
Paper number
353
Field
Machine Learning
arXiv ID
2606.04662

Key points

  • A second-order Taylor expansion decomposes the loss reduction of Muon and Adam. The first-order gains are similar, but the difference appears in the second-order curvature penalty.
  • The update norms are similar, but Muon has lower NDS (normalized direction sharpness), so its curvature cost is smaller.
  • Controlled experiments with Zipf-PCFG show that data imbalance amplifies Muon's NDS advantage.
  • In the middle and later stages of training, Muon's lower NDS comes mainly from reduced within-layer curvature.
  • On heterogeneous-curvature second-order problems, Muon allocates update energy evenly across high- and low-curvature directions.
  • It is rigorously proven that when curvature heterogeneity is sufficiently large, Muon achieves a lower loss than GD.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)