A Survey of Reinforcement Learning for Large Reasoning Models

Published
Source
arXiv
Paper number
085
Field
Reasoning / RL / Survey
arXiv ID
2509.08827

Key points

  • The rapid success of RL in transforming LLMs into LRMs faces substantial challenges related to compute, optimal algorithm design, and scalable high-quality training data.
  • Existing RL applications for LLMs have focused mainly on human alignment, overlooking RL's potential to fundamentally improve complex reasoning abilities.
  • The lack of a unified framework for understanding and navigating the complex, fast-moving landscape of RL for reasoning models hinders systematic research and development.
  • The authors conduct a comprehensive literature review and organize the field into its foundational components: reward design, policy optimization, and sampling strategies.
  • The survey introduces and highlights reinforcement learning with verifiable rewards, or RLVR, as a key paradigm for strengthening reasoning on tasks with objective feedback.
  • It clarifies and discusses foundational questions and ongoing debates, such as RL's role in sharpening versus discovering reasoning and the interaction between RL and supervised fine-tuning, or SFT.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)