Score Centering Stabilizes Off-policy Reinforcement Learning
- Published
- Source
- arXiv
- Paper number
- 1095
- Field
- Machine Learning
- arXiv ID
- 2609.20807
Key points
- Identified the cause of RL instability as 'drift' (engine discrepancies accumulating at every training step) and mathematically derived an additive correction term that cancels it.
- Under low-precision quantization such as FP8/INT4 over 800 training steps, score centering matched or outperformed major correction methods including PPO and DAPO.
- Because it is an additive rather than ratio-based correction, it composes with importance sampling, and the combination beat pure importance sampling in stale rollout experiments.
- Validated with Qwen3-0.6B (Countdown game) and Qwen3-30B (math data), and released the code on GitHub.
Paper links
External research summaries. These are not HDATF publications or measured product results.