Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate
- Published
- Source
- arXiv
- Paper number
- 203
- Field
- Machine Learning
- arXiv ID
- 2605.21486
Key points
- For LayerNorm learning rate, SP uses Θ(1/n), while μP uses Θ(1).
- For attention scaling, SP uses 1/d, and μP uses 1/d.
- The model is trained for a fixed number of iterations regardless of width.
Paper links
External research summaries. These are not HDATF publications or measured product results.