s1: Simple test-time scaling
- Published
- Source
- arXiv
- Paper number
- 025
- Field
- Reasoning / Efficiency
- arXiv ID
- 2501.19393
Key points
- Although test-time scaling has emerged in closed models such as OpenAI's o1, there has been a lack of open-source models that exhibit clear and accessible test-time scaling behavior.
- Traditional large language model performance improvements have depended mainly on compute-intensive training-time scaling, which has made frontier research inaccessible to many.
- Previous open-source attempts to reproduce test-time scaling were often either complex, such as Monte Carlo tree search or multi-agent systems, or still resource-intensive, such as millions of RL samples.
- We curated s1K, a high-quality reasoning dataset of 1,000 samples carefully selected from more than 59,000 examples by quality, difficulty, and diversity.
- We fine-tuned the Qwen2.5-32B-Instruct base model on s1K with supervised fine-tuning, focusing the loss only on reasoning traces and solutions.
- We developed Budget Forcing, a new and simple decoding-time intervention that directly controls thought tokens by forcing generation to stop or by suppressing the end-of-thought token and appending Wait to induce more reasoning.
Paper links
External research summaries. These are not HDATF publications or measured product results.