Learning When to Think: Adaptive Reasoning for Test-Time Compute Allocation
- Published
- Source
- arXiv
- Paper number
- 977
- Field
- AI / General
- arXiv ID
- 2608.20256
Key points
- The model directly selects one of three modes, NoThink, Short, or Long, as the first token of its answer, and this routing is trained jointly end to end within reinforcement learning (GRPO), without a separate router.
- It combines mode-specific token limits of 1,024/3,000/unlimited with a mode-balancing reward to prevent routing collapse, where one mode crowds out the others.
- It showed that the Short mode achieved higher accuracy than Long, demonstrating that the router correctly allocates problems according to difficulty.
- On MATH500, it maintained accuracy of 0.782, compared with the previous 0.796, while reducing average response length by 41%, from 4,796 to 2,811 tokens.
- Without retraining, it saved 76% of tokens on GSM8K, while maintaining equivalent accuracy without reducing the budget on the difficult AIME benchmark, showing that difficulty-based allocation worked as intended.
Paper links
External research summaries. These are not HDATF publications or measured product results.