OPRD: On-Policy Representation Distillation
- Published
- Source
- arXiv
- Paper number
- 350
- Field
- Machine Learning
- arXiv ID
- 2606.06021
Key points
- It moves distillation from output-space matching to hidden-state-space matching, overcoming the structural limitations of all prior OPD variants.
- It removes the variance problem of Monte Carlo KL estimation for large-vocabulary models such as Qwen with about 150K vocabulary by using deterministic MSE instead.
- The teacher's layer-wise and position-wise intermediate representations, which LM-head projection would discard, are transferred directly to the student.
- It fully closes the student-teacher gap on three competitive math benchmarks, AIME 2024, AIME 2025, and AIMO.
- It achieves 1.44x faster training than top-k OPD on the same hardware and uses up to 54 percent less actor-update memory.
- Its memory and distributed advantages are especially strong in expanded scenarios such as multi-model RL merging and on-policy self-distillation, or OPSD.
Paper links
External research summaries. These are not HDATF publications or measured product results.