Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Published
Source
arXiv
Paper number
060
Field
Reasoning / RL
arXiv ID
2504.13837

Key points

  • It is commonly believed that, like traditional RL discovering new strategies, RLVR gives LLMs new reasoning abilities.
  • No rigorous study has determined whether current RLVR truly expands the range of problems LLMs can solve or merely optimizes sampling of existing solutions.
  • Traditional evaluation metrics often fail to capture an LLM's full reasoning potential or its reasoning capability boundary.
  • Pass@k is introduced as a robust metric that defines the LLM's reasoning capability boundary by measuring the fraction of problems solvable within k attempts.
  • The study systematically evaluates reasoning ability across diverse LLM families, sizes, multiple RL algorithms such as PPO and GRPO, and task domains such as math, code, and visual reasoning.
  • Deep analysis methods, including perplexity analysis and solvable-problem coverage, are used to determine whether RLVR-generated reasoning paths are genuinely novel or already lie within the base model's distribution.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)