DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Published
- Source
- arXiv
- Paper number
- 022
- Field
- Reasoning / RL
- arXiv ID
- 2501.12948
Key points
- Existing LLM reasoning methods rely heavily on labor-intensive human-annotated data, creating scalability problems and introducing human cognitive biases.
- Training models to imitate human reasoning paths can cap performance and block the discovery of potentially better non-human strategies.
- Without initial supervised fine-tuning on reasoning tasks, the paper applies DeepSeek-R1-Zero, a pure outcome-based reinforcement learning approach, directly to the strong base LLM DeepSeek-V3-Base.
- It implements a multi-stage training framework, DeepSeek-R1, that combines an initial readability-oriented SFT, continued RL with a language-consistency reward, and a second SFT stage using both reasoning and non-reasoning data.
- It uses knowledge distillation to transfer DeepSeek-R1's advanced reasoning ability to smaller open-source models, improving accessibility.
- The work shows that sophisticated reasoning can emerge in LLMs through outcome-based reinforcement learning without explicit human-annotated reasoning traces.
Paper links
External research summaries. These are not HDATF publications or measured product results.