Scaling Laws for Looped Mixture of Experts

Published
Source
arXiv
Paper number
1162
Field
efficiency
arXiv ID
2609.40316

Key points

  • Unlike prior work that studied recurrence (looping) and sparsity (MoE) separately, it combined both axes into a single law and predicted held-out loss more accurately.
  • It showed with data that the recurrence benefit saturates rather than growing indefinitely, and that having more experts sustains the benefit longer.
  • Sparsity delivers about 3x active-parameter efficiency, recurrence about 2x total-parameter efficiency on reasoning, and combining them pushes the performance frontier further.
  • At trillion-token scale with matched training compute, a looped MoE matched a non-looped MoE twice its size, giving a practical guide for memory and compute budgets.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)