Training-Free Looped Transformers
- Published
- Source
- arXiv
- Paper number
- 230
- Field
- Machine Learning
- arXiv ID
- 2605.23872
Key points
- This method delivered consistent gains on knowledge-intensive benchmarks; for example, Qwen3-4B-Instruct improved by 2.64 percentage points on MMLU-Pro.
- Unlike simple looping, which collapses as more iterations are added, the damped RK method remains stable or keeps improving as K increases, as shown in Figure 3.
- The extra cost is proportional to the number of layers passed through, but because only a small window, such as four layers, is looped, the total runtime overhead is often only about 20 percent for a 32-layer model when the full decode loop is used; the authors also introduce a bypass mode that loops only during initial prompt processing, which makes token-generation overhead close to zero.
Paper links
External research summaries. These are not HDATF publications or measured product results.