A First-Principles Derivation of LLM Policy Optimization: From Expected Reward to GRPO and Its Structural Extensions

Published
Source
arXiv
Paper number
435
Field
AI / General
arXiv ID
2606.16733

Key points

  • It unifies all policy gradient methods into two axes, trajectory side and reward side, for J(theta) = E[R(tau)].
  • It explains the evolution from REINFORCE to PPO to GRPO to post-GRPO variants through the failure modes and interventions at each stage.
  • It extends the same two-axis view to agentic RL and hybrid GRPO-OPD.
  • It identifies coupled failures that cannot be solved by changing only one axis and points to directions for next-generation algorithm design.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)