Scaling Muon for Diffusion Transformers

Published
Source
arXiv
Paper number
984
Field
Machine Learning
arXiv ID
2608.20818

Key points

  • It found a tendency for Muon to improve generation-quality metrics by 12.9–19.1% over AdamW in diffusion transformers with 1.3 billion–15 billion parameters.
  • It proposed Periodic Row-wise Muon, which performs a full spectral update only every three steps and replaces the remaining updates with inexpensive row-wise updates.
  • Optimizer time fell by 46.9–54.3%, total training-step time by 15.7–24.3%, and logical communication volume by 66.7%.
  • On the 9-billion-parameter model, generation quality improved by a further 4.5% over the original Muon, demonstrating both cost reduction and quality gains.
  • The experiments were limited to one diffusion-model family, one dataset and resolution, and a 32-node H100 configuration, and full momentum communication remains on update steps.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)