Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- Published
- Source
- arXiv
- Paper number
- 060
- Field
- Reasoning / RL
- arXiv ID
- 2504.13837
Key points
- It is commonly believed that, like traditional RL discovering new strategies, RLVR gives LLMs new reasoning abilities.
- No rigorous study has determined whether current RLVR truly expands the range of problems LLMs can solve or merely optimizes sampling of existing solutions.
- Traditional evaluation metrics often fail to capture an LLM's full reasoning potential or its reasoning capability boundary.
- Pass@k is introduced as a robust metric that defines the LLM's reasoning capability boundary by measuring the fraction of problems solvable within k attempts.
- The study systematically evaluates reasoning ability across diverse LLM families, sizes, multiple RL algorithms such as PPO and GRPO, and task domains such as math, code, and visual reasoning.
- Deep analysis methods, including perplexity analysis and solvable-problem coverage, are used to determine whether RLVR-generated reasoning paths are genuinely novel or already lie within the base model's distribution.
Paper links
External research summaries. These are not HDATF publications or measured product results.