Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors

Published
Source
arXiv
Paper number
502
Field
Machine Learning
arXiv ID
2606.25971

Key points

  • It mathematically analyzes interference between weight magnitude and direction: direction updates are inversely proportional to the current magnitude, and magnitude grows as a byproduct of direction learning.
  • It decomposes weights inside the optimizer into fixed-norm directions plus learnable gains, so it can be applied without changing the model architecture.
  • Weight decay and warmup become unnecessary, greatly simplifying large-scale training recipes.
  • The optimal learning rate transfers across model widths, so no LR retuning is needed, similar to µP but obtained directly.
  • On large MoE models, MuonMD reaches the same loss as AdamW with about half the compute.
  • Fixing the direction on the sphere captures most of the improvement, and learnable gains provide additional accuracy gains.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)