How to Train a Critic Stably and Efficiently

Published
Source
arXiv
Paper number
1000
Field
Machine Learning
arXiv ID
2608.23566

Key points

  • Controlled experiments isolated the causes of critic instability one by one (unbounded value predictions, batch-level normalization, and others).
  • It proposed the BPCO recipe, combining DPPO + reward-range constraints + Monte Carlo targets + unnormalized advantages.
  • Sampling just 1 response per prompt matched or exceeded group-based (GRPO) baselines.
  • Because the critic is used only during training, information that must be hidden from the policy, such as correct answers or scoring rubrics, can be provided only to the critic. This allows reinforcement learning to extend to tasks rewarded through rubrics.
  • However, the paper states that its evidence is limited to mathematics and rubric-based rewards, that the reward range must be known in advance, and that the comparisons do not account for the additional computation and memory used to train the critic.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)