DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Published
Source
arXiv
Paper number
022
Field
Reasoning / RL
arXiv ID
2501.12948

Key points

  • Existing LLM reasoning methods rely heavily on labor-intensive human-annotated data, creating scalability problems and introducing human cognitive biases.
  • Training models to imitate human reasoning paths can cap performance and block the discovery of potentially better non-human strategies.
  • Without initial supervised fine-tuning on reasoning tasks, the paper applies DeepSeek-R1-Zero, a pure outcome-based reinforcement learning approach, directly to the strong base LLM DeepSeek-V3-Base.
  • It implements a multi-stage training framework, DeepSeek-R1, that combines an initial readability-oriented SFT, continued RL with a language-consistency reward, and a second SFT stage using both reasoning and non-reasoning data.
  • It uses knowledge distillation to transfer DeepSeek-R1's advanced reasoning ability to smaller open-source models, improving accessibility.
  • The work shows that sophisticated reasoning can emerge in LLMs through outcome-based reinforcement learning without explicit human-annotated reasoning traces.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)