OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation

Published
Source
arXiv
Paper number
349
Field
Machine Learning
arXiv ID
2606.06096

Key points

  • We propose an unbiased gradient estimator for order-statistic objective functions, and by changing only the rank weights, it can express many distributional objectives such as VaR, CVaR, trimmed mean, and best-of-K.
  • It is a plug-and-play design that can be applied to existing likelihood-ratio and reparameterization updates in one line of code by reformulating them as reward transformations.
  • Top-M@K shows consistent gains over the existing Max@K, which is the special case max@K with m=1, on both pass@1 and pass@256.
  • Combining a correct-answer reward for Top-M and a length penalty for Bottom-M shortens responses far more effectively than the simple scalarization used by GRPO.
  • Choosing K, the target size, and m, the number of top ranks, controls the exploration-exploitation tradeoff, with larger K improving pass@256 while reducing pass@1.
  • It has effectively negligible overhead, with an O(N log N) sorting cost.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)