Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
- Published
- Source
- arXiv
- Paper number
- 1053
- Field
- Machine Learning
- arXiv ID
- 2608.31046
Key points
- Noise in OPD's teacher signals, penalizing correct answers and rewarding incorrect ones, reached 30.6% for a 4B teacher and 50.6% for a 235B teacher, yet training after removing this noise produced essentially the same results.
- It showed that OPD's gains concentrate on low-probability tokens and that a single fixed negative advantage can replace the teacher signal.
- OPSA, with no external supervision at all, raised AIME24 Avg@32 by +35.41 points, a relative gain of 263%, and exceeded OPD by 16.77 points.
- It at least doubled Pass@32 on all 3 mathematical benchmarks: AIME24/25 and HMMT.
- At high-entropy branching points where probabilities diverge, it distributes probability evenly among candidate tokens, preserving exploration diversity while sharpening the distribution.
Paper links
External research summaries. These are not HDATF publications or measured product results.