On the Overthinking of LLMs

Published
Source
arXiv
Paper number
004
Field
Reasoning / Efficiency
arXiv ID
2412.21187

Key points

  • The problem is that o1-family models often spend long reasoning traces on simple tasks, which brings only limited accuracy gains while increasing latency and cost.
  • In evaluation, the paper proposes outcome-level and process-level efficiency metrics to measure whether reasoning compute is being used sensibly.
  • The method trains the model with a self-learning strategy to simplify reasoning while preserving answer quality across tasks of different difficulty.
  • The implication is that future reasoning systems need adaptive test-time compute policies, not a habit of using verbose chains of thought for every query.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)