Concise Reasoning via Reinforcement Learning

Published
Source
arXiv
Paper number
055
Field
Reasoning / Efficiency
arXiv ID
2504.05185

Key points

  • Large language models strengthened for reasoning tasks through reinforcement learning, or RL, tend to produce excessively verbose chain-of-thought outputs, which increases compute cost and inference latency.
  • The prior literature contains conflicting observations about whether response length is necessary for reasoning accuracy, with some studies finding that longer responses improve accuracy while others observe diminishing returns.
  • The fundamental mechanism by which RL training affects response length, and whether this verbosity is an essential part of reasoning or merely a byproduct, is not well understood.
  • The paper casts reasoning as a Markov decision process and performs a theoretical analysis of the loss dynamics of Proximal Policy Optimization, or PPO, and Generalized Reinforcement Learning with Policy Optimization, or GRPO, to explain how negative rewards push the policy toward verbosity.
  • It proposes a new two-stage RL training procedure that first strengthens reasoning on hard problems and then explicitly enforces conciseness on a small dataset of problems that are sometimes solvable.
  • It uses a built-in property of PPO, with generalized advantage estimation lambda below 1, where positive rewards encourage shorter responses and negative rewards encourage longer ones, to guide the model toward conciseness without sacrificing accuracy.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)