Weak-to-Strong Generalization via Direct On-Policy Distillation
- Published
- Source
- arXiv
- Paper number
- 577
- Field
- Machine Learning
- arXiv ID
- 2607.05394
Key points
- The key idea is that the log ratio between a post-RL weak model and a pre-RL reference is equal to the implicit reward of KL-regularized RL.
- The RL results of a weak 1.5B model transfer to stronger models such as Qwen3-1.7B/4B and R1-Distill-7B, improving all of them.
- Qwen3-1.7B's AIME 2024 accuracy rises from 48.3% to 62.4% with 8x A100s over 4 hours, matching direct RL with Polaris, which uses 32x A100s for a week.
- By contrast, simply imitating the final policy of the weak model, that is, ordinary OPD, actually hurts the stronger student's performance.
- Response length and the KL coefficient are the key control variables for transfer reliability, and adaptive KL stabilizes average reward.
- Consistent improvements are observed across several teacher pairs, including R1-Distill to JustRL and Nemotron to QuestA.
Paper links
External research summaries. These are not HDATF publications or measured product results.