Score Centering Stabilizes Off-policy Reinforcement Learning

Published
Source
arXiv
Paper number
1095
Field
Machine Learning
arXiv ID
2609.20807

Key points

  • Identified the cause of RL instability as 'drift' (engine discrepancies accumulating at every training step) and mathematically derived an additive correction term that cancels it.
  • Under low-precision quantization such as FP8/INT4 over 800 training steps, score centering matched or outperformed major correction methods including PPO and DAPO.
  • Because it is an additive rather than ratio-based correction, it composes with importance sampling, and the combination beat pure importance sampling in stale rollout experiments.
  • Validated with Qwen3-0.6B (Countdown game) and Qwen3-30B (math data), and released the code on GitHub.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)