Z1: Efficient Test-time Scaling with Code

Published
Source
arXiv
Paper number
052
Field
Reasoning / Efficiency
arXiv ID
2504.00810

Key points

  • Large language models consume too many tokens during reasoning tasks, even on simple problems that do not need elaborate thinking steps.
  • Prior approaches lack an efficient way to scale test-time reasoning while controlling token usage.
  • The model is fine-tuned on a carefully selected code reasoning dataset of 107K examples with balanced short and long trajectories.
  • The method implements a Shifted Thinking Window that limits reasoning tokens and forces a direct answer once the threshold is exceeded.
  • It removes context separators so that reasoning depth can be adjusted more flexibly.
  • The results suggest that longer reasoning trajectories and larger dataset scale are key to eliciting effective reasoning.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)