Dion3: Full-Stack Orthogonal Updates
- Published
- Source
- arXiv
- Paper number
- 891
- Field
- Machine Learning
- arXiv ID
- 2608.11612
Key points
- Gram Newton-Schulz rewrites existing methods into mathematically equivalent form and reduces compute by iterating on small symmetric matrices instead of large rectangular ones.
- It adds a symmetric GPU kernel and a grouped-communication strategy for matrices of the same shape. In the 8-way setup for a 1-billion-parameter model, grouped communication alone cuts optimizer-step time from 80.7 milliseconds to 52.1 milliseconds.
- At each step, it orthogonalizes only a subset of rows in the momentum matrix to reduce compute and communication, and it improves both speed and loss over the previous compression method, Dion.
- In experiments training 3B to 14B models on 10 billion tokens, all models achieved lower validation loss, and the 14B model had 0.027 lower loss than NorMuon and 0.7 percentage points higher average accuracy.
- The reported up to 6x speedup applies only to the optimizer step, not full training, and the authors analyze that this step accounts for roughly 1 percent to 17 percent of total training time depending on the setup.
Paper links
External research summaries. These are not HDATF publications or measured product results.