s1: Simple test-time scaling

Published
Source
arXiv
Paper number
025
Field
Reasoning / Efficiency
arXiv ID
2501.19393

Key points

  • Although test-time scaling has emerged in closed models such as OpenAI's o1, there has been a lack of open-source models that exhibit clear and accessible test-time scaling behavior.
  • Traditional large language model performance improvements have depended mainly on compute-intensive training-time scaling, which has made frontier research inaccessible to many.
  • Previous open-source attempts to reproduce test-time scaling were often either complex, such as Monte Carlo tree search or multi-agent systems, or still resource-intensive, such as millions of RL samples.
  • We curated s1K, a high-quality reasoning dataset of 1,000 samples carefully selected from more than 59,000 examples by quality, difficulty, and diversity.
  • We fine-tuned the Qwen2.5-32B-Instruct base model on s1K with supervised fine-tuning, focusing the loss only on reasoning traces and solutions.
  • We developed Budget Forcing, a new and simple decoding-time intervention that directly controls thought tokens by forcing generation to stop or by suppressing the end-of-thought token and appending Wait to induce more reasoning.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)