Improving Test-Time Scaling with Adaptive Looped Transformers

Published
Source
arXiv
Paper number
1136
Field
LLMs / NLP
arXiv ID
2609.35748

Key points

  • Co-post-trained an iteration decider with the backbone using online labels of whether one more loop iteration improves the prediction.
  • Improved the accuracy-per-compute-doubling slope of test-time scaling by 53% over the non-looped baseline (2.74 vs 1.79).
  • Exceeded the baseline's peak AIME accuracy by about 3.4 points at matched compute, with gains growing from +2.8 at depth 2 to +3.9 at depth 8.
  • Focused extra computation only on tokens that benefit from looping, unlike fixed-depth looping which iterates on every token.
  • Achieved these results purely through post-training of a 1.7B model.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)