Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining

Published
Source
arXiv
Paper number
1012
Field
Machine Learning
arXiv ID
2608.24814

Key points

  • It proposed a new perspective: first design the schedule of ELR, the ratio between learning rate and parameter magnitude, rather than tuning the two separately.
  • Matching ELR alone produced nearly identical loss curves across different optimizers, architectures, data, and model sizes (with an average error on the order of 10^-3).
  • The effects of weight decay and Hyperball (a technique that constrains parameters to a specified magnitude) on loss were also explained by the ELR schedules each produces.
  • ELR-based scaling laws remained predictive across different magnitude-control methods. Using learning rate instead increased prediction error 12.9-fold.
  • Condition-specific experiments also identified when the law holds precisely (normalization design and rate of change).

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)