Scaling Muon for Diffusion Transformers
- Published
- Source
- arXiv
- Paper number
- 984
- Field
- Machine Learning
- arXiv ID
- 2608.20818
Key points
- It found a tendency for Muon to improve generation-quality metrics by 12.9–19.1% over AdamW in diffusion transformers with 1.3 billion–15 billion parameters.
- It proposed Periodic Row-wise Muon, which performs a full spectral update only every three steps and replaces the remaining updates with inexpensive row-wise updates.
- Optimizer time fell by 46.9–54.3%, total training-step time by 15.7–24.3%, and logical communication volume by 66.7%.
- On the 9-billion-parameter model, generation quality improved by a further 4.5% over the original Muon, demonstrating both cost reduction and quality gains.
- The experiments were limited to one diffusion-model family, one dataset and resolution, and a 32-node H100 configuration, and full momentum communication remains on update steps.
Paper links
External research summaries. These are not HDATF publications or measured product results.