Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning

Published
Source
arXiv
Paper number
807
Field
AI / General
arXiv ID
2608.01418

Key points

  • PNPO multiplies likelihood ratios through each token position and then takes their geometric mean, retaining causal prefix information while narrowing the weight range.
  • Unlike GSPO, which assigns one weight to an entire response, it uses a different prefix average at each position, but it is a biased approximation that gives up exact state-action distribution correction.
  • Under fourfold reuse, its peak average across three mathematical benchmarks was 50.24, 3.00 points above GSPO, but the difference was inconsistent under single-use training.
  • Under the same budget of 2400 updates, using only 150 fresh rollout batches produced results similar to using 600 batches for single-use training, demonstrating the potential of rollout reuse.
  • The study used only one model, three mathematical tasks, and one run per setting, without separately ablating components, so the effect of prefix normalization alone and its effects on asynchronous and replay data have not yet been isolated.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)