OPRD: On-Policy Representation Distillation

Published
Source
arXiv
Paper number
350
Field
Machine Learning
arXiv ID
2606.06021

Key points

  • It moves distillation from output-space matching to hidden-state-space matching, overcoming the structural limitations of all prior OPD variants.
  • It removes the variance problem of Monte Carlo KL estimation for large-vocabulary models such as Qwen with about 150K vocabulary by using deterministic MSE instead.
  • The teacher's layer-wise and position-wise intermediate representations, which LM-head projection would discard, are transferred directly to the student.
  • It fully closes the student-teacher gap on three competitive math benchmarks, AIME 2024, AIME 2025, and AIMO.
  • It achieves 1.44x faster training than top-k OPD on the same hardware and uses up to 54 percent less actor-update memory.
  • Its memory and distributed advantages are especially strong in expanded scenarios such as multi-model RL merging and on-policy self-distillation, or OPSD.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)