Dion3: Full-Stack Orthogonal Updates

Published
Source
arXiv
Paper number
891
Field
Machine Learning
arXiv ID
2608.11612

Key points

  • Gram Newton-Schulz rewrites existing methods into mathematically equivalent form and reduces compute by iterating on small symmetric matrices instead of large rectangular ones.
  • It adds a symmetric GPU kernel and a grouped-communication strategy for matrices of the same shape. In the 8-way setup for a 1-billion-parameter model, grouped communication alone cuts optimizer-step time from 80.7 milliseconds to 52.1 milliseconds.
  • At each step, it orthogonalizes only a subset of rows in the momentum matrix to reduce compute and communication, and it improves both speed and loss over the previous compression method, Dion.
  • In experiments training 3B to 14B models on 10 billion tokens, all models achieved lower validation loss, and the 14B model had 0.027 lower loss than NorMuon and 0.7 percentage points higher average accuracy.
  • The reported up to 6x speedup applies only to the optimizer step, not full training, and the authors analyze that this step accounts for roughly 1 percent to 17 percent of total training time depending on the setup.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)