Improving Test-Time Scaling with Adaptive Looped Transformers
- Published
- Source
- arXiv
- Paper number
- 1136
- Field
- LLMs / NLP
- arXiv ID
- 2609.35748
Key points
- Co-post-trained an iteration decider with the backbone using online labels of whether one more loop iteration improves the prediction.
- Improved the accuracy-per-compute-doubling slope of test-time scaling by 53% over the non-looped baseline (2.74 vs 1.79).
- Exceeded the baseline's peak AIME accuracy by about 3.4 points at matched compute, with gains growing from +2.8 at depth 2 to +3.9 at depth 8.
- Focused extra computation only on tokens that benefit from looping, unlike fixed-depth looping which iterates on every token.
- Achieved these results purely through post-training of a 1.7B model.
Paper links
External research summaries. These are not HDATF publications or measured product results.