Prefix Sliding for efficient test-time scaling
- Published
- Source
- arXiv
- Paper number
- 1028
- Field
- LLMs / NLP
- arXiv ID
- 2608.26070
Key points
- Attention analysis showed that intermediate reasoning tokens quickly lose importance, leaving only the prefix and recent tokens important.
- Applied to an existing model (Qwen3) without training, it preserved performance while making inference 3 times faster.
- Combined with RL training, it enabled training on reasoning sequences longer than 100,000 tokens, addressing the problem of discarding long rollouts.
- It confirmed advantages in both performance and speed over alternatives such as summarizing intermediate tokens or using a simple sliding window.
- It also identified a limitation: tasks that require remembering early code for a long time, such as the coding benchmark LiveCodeBench, need a large window (16384).
Paper links
External research summaries. These are not HDATF publications or measured product results.