NaturalThoughts: Selecting and Distilling Reasoning Traces for General Reasoning Tasks

Published
Source
arXiv
Paper number
071
Field
Reasoning / Data
arXiv ID
2507.01921

Key points

  • There was no systematic understanding of which characteristics of reasoning demonstrations are most effective for distilling reasoning ability into smaller language models.
  • The prevailing Less is More hypothesis, which claims that a few handpicked examples are enough for reasoning distillation, mostly came from domain-specific settings and needs reevaluation for general reasoning.
  • Large language models often overthink and produce excessively long reasoning traces, which increases inference cost and reduces deployment efficiency in practice.
  • The authors built NaturalThoughts, a dataset that annotates reasoning traces generated by DeepSeek-R1, the teacher, for diverse NaturalReasoning questions by domain, meta-reasoning strategy, and verbosity.
  • They systematically explored data selection strategies based on diversity, such as question topic, semantic embedding, and reasoning strategy, as well as difficulty, such as answer length, verbosity, and model agreement.
  • They proposed and evaluated a mixed System 1/System 2 distillation approach, including random and difficulty-based mixtures, so student models can adjust reasoning depth at inference time.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)