Prefix Sliding for efficient test-time scaling

Published
Source
arXiv
Paper number
1028
Field
LLMs / NLP
arXiv ID
2608.26070

Key points

  • Attention analysis showed that intermediate reasoning tokens quickly lose importance, leaving only the prefix and recent tokens important.
  • Applied to an existing model (Qwen3) without training, it preserved performance while making inference 3 times faster.
  • Combined with RL training, it enabled training on reasoning sequences longer than 100,000 tokens, addressing the problem of discarding long rollouts.
  • It confirmed advantages in both performance and speed over alternatives such as summarizing intermediate tokens or using a simple sliding window.
  • It also identified a limitation: tasks that require remembering early code for a long time, such as the coding benchmark LiveCodeBench, need a large window (16384).

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)