The Loss Does Not See the Basis, but Adam Does
- Published
- Source
- arXiv
- Paper number
- 837
- Field
- Machine Learning
- arXiv ID
- 2608.05136
Key points
- We proved that gauge equivalence, or rotational invariance, is necessary to preserve low-rank bias.
- A classification of nine optimizers shows that GD, Muon, and Shampoo are equivalent, while Adam, RMSProp, and Lion are not.
- In matrix sensing, equivalence classes achieve recovery error from 0.00 to 0.29, whereas coordinate-wise classes are at 0.42 to 0.57.
- In Transformers, Adam separates gauge-equivalent initializations from the first step and creates a 56% difference.
- Interpolating from coordinate-wise forms to equivalent scalars improves monotonically, proving that heterogeneity is the cause.
Paper links
External research summaries. These are not HDATF publications or measured product results.