Scaling Laws for Looped Mixture of Experts
- Published
- Source
- arXiv
- Paper number
- 1162
- Field
- efficiency
- arXiv ID
- 2609.40316
Key points
- Unlike prior work that studied recurrence (looping) and sparsity (MoE) separately, it combined both axes into a single law and predicted held-out loss more accurately.
- It showed with data that the recurrence benefit saturates rather than growing indefinitely, and that having more experts sustains the benefit longer.
- Sparsity delivers about 3x active-parameter efficiency, recurrence about 2x total-parameter efficiency on reasoning, and combining them pushes the performance frontier further.
- At trillion-token scale with matched training compute, a looped MoE matched a non-looped MoE twice its size, giving a practical guide for memory and compute budgets.
Paper links
External research summaries. These are not HDATF publications or measured product results.