The Loss Does Not See the Basis, but Adam Does

Published
Source
arXiv
Paper number
837
Field
Machine Learning
arXiv ID
2608.05136

Key points

  • We proved that gauge equivalence, or rotational invariance, is necessary to preserve low-rank bias.
  • A classification of nine optimizers shows that GD, Muon, and Shampoo are equivalent, while Adam, RMSProp, and Lion are not.
  • In matrix sensing, equivalence classes achieve recovery error from 0.00 to 0.29, whereas coordinate-wise classes are at 0.42 to 0.57.
  • In Transformers, Adam separates gauge-equivalent initializations from the first step and creates a 56% difference.
  • Interpolating from coordinate-wise forms to equivalent scalars improves monotonically, proving that heterogeneity is the cause.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)