Weak-to-Strong Generalization via Direct On-Policy Distillation

Published
Source
arXiv
Paper number
577
Field
Machine Learning
arXiv ID
2607.05394

Key points

  • The key idea is that the log ratio between a post-RL weak model and a pre-RL reference is equal to the implicit reward of KL-regularized RL.
  • The RL results of a weak 1.5B model transfer to stronger models such as Qwen3-1.7B/4B and R1-Distill-7B, improving all of them.
  • Qwen3-1.7B's AIME 2024 accuracy rises from 48.3% to 62.4% with 8x A100s over 4 hours, matching direct RL with Polaris, which uses 32x A100s for a week.
  • By contrast, simply imitating the final policy of the weak model, that is, ordinary OPD, actually hurts the stronger student's performance.
  • Response length and the KL coefficient are the key control variables for transfer reliability, and adaptive KL stabilizes average reward.
  • Consistent improvements are observed across several teacher pairs, including R1-Distill to JustRL and Nemotron to QuestA.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)