NaturalThoughts: Selecting and Distilling Reasoning Traces for General Reasoning Tasks
- Published
- Source
- arXiv
- Paper number
- 071
- Field
- Reasoning / Data
- arXiv ID
- 2507.01921
Key points
- There was no systematic understanding of which characteristics of reasoning demonstrations are most effective for distilling reasoning ability into smaller language models.
- The prevailing Less is More hypothesis, which claims that a few handpicked examples are enough for reasoning distillation, mostly came from domain-specific settings and needs reevaluation for general reasoning.
- Large language models often overthink and produce excessively long reasoning traces, which increases inference cost and reduces deployment efficiency in practice.
- The authors built NaturalThoughts, a dataset that annotates reasoning traces generated by DeepSeek-R1, the teacher, for diverse NaturalReasoning questions by domain, meta-reasoning strategy, and verbosity.
- They systematically explored data selection strategies based on diversity, such as question topic, semantic embedding, and reasoning strategy, as well as difficulty, such as answer length, verbosity, and model agreement.
- They proposed and evaluated a mixed System 1/System 2 distillation approach, including random and difficulty-based mixtures, so student models can adjust reasoning depth at inference time.
Paper links
External research summaries. These are not HDATF publications or measured product results.