Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts
- Published
- Source
- arXiv
- Paper number
- 970
- Field
- Machine Learning
- arXiv ID
- 2608.20061
Key points
- In stage 1, it formalized μP for MoE models using MLA attention and the Muon optimizer, and showed that the optimal learning rate found on a base proxy model (0.6B total, 0.3B active) transfers unchanged to models with 2, 4, and 8 times the width (up to 30.7B total, 3.6B active).
- In stage 2, it stops training during the stable phase without a decay phase, extracts multiple checkpoints using EMA (exponential moving average) weights, fits optimal learning rates over 255B to 502B tokens using log-log linear regression, and extrapolates to longer horizons.
- The regression fit was R²=0.95, and it predicted 3.85×10⁻⁴ as the optimal learning rate for training on 10 trillion tokens.
- Using the predicted learning rate, it trained a foundation model with 155B total and 17B active parameters from scratch on 10 trillion tokens. The target training run required roughly 98 times the total computation of all proxy experiments, so avoiding a search saves substantial cost.
- However, the authors state that validation is limited to the MLA and Muon combination, batch size was deliberately left out, increasing sparsity together with width prevented isolation of the sparsity axis's effect, and an exhaustive search to verify whether the predicted learning rate is truly optimal is computationally infeasible.
Paper links
External research summaries. These are not HDATF publications or measured product results.