Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate

Published
Source
arXiv
Paper number
203
Field
Machine Learning
arXiv ID
2605.21486

Key points

  • For LayerNorm learning rate, SP uses Θ(1/n), while μP uses Θ(1).
  • For attention scaling, SP uses 1/d, and μP uses 1/d.
  • The model is trained for a fixed number of iterations regardless of width.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)