Z1: Efficient Test-time Scaling with Code
- Published
- Source
- arXiv
- Paper number
- 052
- Field
- Reasoning / Efficiency
- arXiv ID
- 2504.00810
Key points
- Large language models consume too many tokens during reasoning tasks, even on simple problems that do not need elaborate thinking steps.
- Prior approaches lack an efficient way to scale test-time reasoning while controlling token usage.
- The model is fine-tuned on a carefully selected code reasoning dataset of 107K examples with balanced short and long trajectories.
- The method implements a Shifted Thinking Window that limits reasoning tokens and forces a direct answer once the threshold is exceeded.
- It removes context separators so that reasoning depth can be adjusted more flexibly.
- The results suggest that longer reasoning trajectories and larger dataset scale are key to eliciting effective reasoning.
Paper links
External research summaries. These are not HDATF publications or measured product results.