Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors
- Published
- Source
- arXiv
- Paper number
- 502
- Field
- Machine Learning
- arXiv ID
- 2606.25971
Key points
- It mathematically analyzes interference between weight magnitude and direction: direction updates are inversely proportional to the current magnitude, and magnitude grows as a byproduct of direction learning.
- It decomposes weights inside the optimizer into fixed-norm directions plus learnable gains, so it can be applied without changing the model architecture.
- Weight decay and warmup become unnecessary, greatly simplifying large-scale training recipes.
- The optimal learning rate transfers across model widths, so no LR retuning is needed, similar to µP but obtained directly.
- On large MoE models, MuonMD reaches the same loss as AdamW with about half the compute.
- Fixing the direction on the sphere captures most of the improvement, and learnable gains provide additional accuracy gains.
Paper links
External research summaries. These are not HDATF publications or measured product results.