Demystifying Long Chain-of-Thought Reasoning in LLMs
- Published
- Source
- arXiv
- Paper number
- 029
- Field
- Reasoning
- arXiv ID
- 2502.03373
Key points
- Large language models still struggle with complex reasoning domains such as advanced math, science QA, and software engineering, even when Chain-of-Thought prompting is used.
- In particular, our understanding of how long CoT capabilities emerge in LLMs remains limited in the context of reinforcement learning training.
- RL training for CoT often runs into unstable growth in CoT length, outputs that exceed context-window limits, and reward hacking unless the design choices are appropriate.
- The researchers systematically studied long CoT reasoning by combining supervised fine-tuning (SFT) and reinforcement learning (RL) with models such as Llama-3.1-8B and Qwen2.5-7B-Math.
- They explored distilling long CoT patterns from a strong emergent model, QwQ-32B-Preview, to curate high-quality SFT data and filtering noisy web-extracted data to expand verifiable reward signals.
- For RL, they introduced a new reward function called Cosine Length-Scaling Reward with an N-gram repetition penalty to stabilize CoT length and control generation behavior during training.
Paper links
External research summaries. These are not HDATF publications or measured product results.