SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

Published
Source
arXiv
Paper number
1063
Field
Machine Learning
arXiv ID
2609.01343

Key points

  • This is the first multi-scale comparison in looping research to simultaneously match all three budgets, compute per token, total parameters, and KV cache, ruling out the objection that the gains merely come from using more computation.
  • The design loops only the middle half of the layers twice, reduces the hidden dimension, and then increases the number of experts to restore knowledge capacity, exploiting a budget trade-off possible only with MoE.
  • It consistently achieved lower loss across 4 scales up to 54B non-embedding parameters and multiple sparsity levels, with scaling-law fits indicating that reaching the same performance requires 6.8–18.0% less training compute.
  • The largest gains were in the code domain, and the gap widened with longer inputs and more examples in the prompt, showing that looping becomes more valuable for structured inputs.
  • On the second pass, attention sinks disappeared and attention shifted toward evidence tokens for the correct answer. In a bracket-matching task, attention concentrated on the first token fell from 0.60 to 0.02, while attention to evidence tokens rose from 0.24 to 0.85.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)