A Survey of Reinforcement Learning for Large Reasoning Models
- Published
- Source
- arXiv
- Paper number
- 085
- Field
- Reasoning / RL / Survey
- arXiv ID
- 2509.08827
Key points
- The rapid success of RL in transforming LLMs into LRMs faces substantial challenges related to compute, optimal algorithm design, and scalable high-quality training data.
- Existing RL applications for LLMs have focused mainly on human alignment, overlooking RL's potential to fundamentally improve complex reasoning abilities.
- The lack of a unified framework for understanding and navigating the complex, fast-moving landscape of RL for reasoning models hinders systematic research and development.
- The authors conduct a comprehensive literature review and organize the field into its foundational components: reward design, policy optimization, and sampling strategies.
- The survey introduces and highlights reinforcement learning with verifiable rewards, or RLVR, as a key paradigm for strengthening reasoning on tasks with objective feedback.
- It clarifies and discusses foundational questions and ongoing debates, such as RL's role in sharpening versus discovering reasoning and the interaction between RL and supervised fine-tuning, or SFT.
Paper links
External research summaries. These are not HDATF publications or measured product results.