On Advantage Estimates for Max@K Policy Gradients

Published
Source
arXiv
Paper number
346
Field
Machine Learning
arXiv ID
2606.06080

Key points

  • The paper mathematically proves that the existing PKPO advantage estimator is unbiased for policy gradients but not centered, and addresses this with a Leave-Two-Out (L2O) baseline.
  • It derives the canonical advantage u_i - v_i for finite-batch max@K and unifies existing pass@K and max@K estimators under a single framework.
  • It is efficient with a vectorized implementation that has O(B^2) time complexity and integrates naturally into group-based RL pipelines.
  • On Llama-3.2-3B-Instruct, it reduces gradient variance by 77.4 percent relative to PKPO and improves pass@256 by 5.2 percent on Qwen2.5-Math-7B.
  • It shows consistent gains over strong post-training baselines for LLMs, including GRPO and Entropy-Adv.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)