How to Train a Critic Stably and Efficiently
- Published
- Source
- arXiv
- Paper number
- 1000
- Field
- Machine Learning
- arXiv ID
- 2608.23566
Key points
- Controlled experiments isolated the causes of critic instability one by one (unbounded value predictions, batch-level normalization, and others).
- It proposed the BPCO recipe, combining DPPO + reward-range constraints + Monte Carlo targets + unnormalized advantages.
- Sampling just 1 response per prompt matched or exceeded group-based (GRPO) baselines.
- Because the critic is used only during training, information that must be hidden from the policy, such as correct answers or scoring rubrics, can be provided only to the critic. This allows reinforcement learning to extend to tasks rewarded through rubrics.
- However, the paper states that its evidence is limited to mathematics and rubric-based rewards, that the reward range must be known in advance, and that the comparisons do not account for the additional computation and memory used to train the critic.
Paper links
External research summaries. These are not HDATF publications or measured product results.