Training-Free Looped Transformers

Published
Source
arXiv
Paper number
230
Field
Machine Learning
arXiv ID
2605.23872

Key points

  • This method delivered consistent gains on knowledge-intensive benchmarks; for example, Qwen3-4B-Instruct improved by 2.64 percentage points on MMLU-Pro.
  • Unlike simple looping, which collapses as more iterations are added, the damped RK method remains stable or keeps improving as K increases, as shown in Figure 3.
  • The extra cost is proportional to the number of layers passed through, but because only a small window, such as four layers, is looped, the total runtime overhead is often only about 20 percent for a 32-layer model when the full decode loop is used; the authors also introduce a bypass mode that loops only during initial prompt processing, which makes token-generation overhead close to zero.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)