Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining
- Published
- Source
- arXiv
- Paper number
- 1012
- Field
- Machine Learning
- arXiv ID
- 2608.24814
Key points
- It proposed a new perspective: first design the schedule of ELR, the ratio between learning rate and parameter magnitude, rather than tuning the two separately.
- Matching ELR alone produced nearly identical loss curves across different optimizers, architectures, data, and model sizes (with an average error on the order of 10^-3).
- The effects of weight decay and Hyperball (a technique that constrains parameters to a specified magnitude) on loss were also explained by the ELR schedules each produces.
- ELR-based scaling laws remained predictive across different magnitude-control methods. Using learning rate instead increased prediction error 12.9-fold.
- Condition-specific experiments also identified when the law holds precisely (normalization design and rate of change).
Paper links
External research summaries. These are not HDATF publications or measured product results.