Demystifying Long Chain-of-Thought Reasoning in LLMs

Published
Source
arXiv
Paper number
029
Field
Reasoning
arXiv ID
2502.03373

Key points

  • Large language models still struggle with complex reasoning domains such as advanced math, science QA, and software engineering, even when Chain-of-Thought prompting is used.
  • In particular, our understanding of how long CoT capabilities emerge in LLMs remains limited in the context of reinforcement learning training.
  • RL training for CoT often runs into unstable growth in CoT length, outputs that exceed context-window limits, and reward hacking unless the design choices are appropriate.
  • The researchers systematically studied long CoT reasoning by combining supervised fine-tuning (SFT) and reinforcement learning (RL) with models such as Llama-3.1-8B and Qwen2.5-7B-Math.
  • They explored distilling long CoT patterns from a strong emergent model, QwQ-32B-Preview, to curate high-quality SFT data and filtering noisy web-extracted data to expand verifiable reward signals.
  • For RL, they introduced a new reward function called Cosine Length-Scaling Reward with an N-gram repetition penalty to stabilize CoT length and control generation behavior during training.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)