Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

Published
Source
arXiv
Paper number
1053
Field
Machine Learning
arXiv ID
2608.31046

Key points

  • Noise in OPD's teacher signals, penalizing correct answers and rewarding incorrect ones, reached 30.6% for a 4B teacher and 50.6% for a 235B teacher, yet training after removing this noise produced essentially the same results.
  • It showed that OPD's gains concentrate on low-probability tokens and that a single fixed negative advantage can replace the teacher signal.
  • OPSA, with no external supervision at all, raised AIME24 Avg@32 by +35.41 points, a relative gain of 263%, and exceeded OPD by 16.77 points.
  • It at least doubled Pass@32 on all 3 mathematical benchmarks: AIME24/25 and HMMT.
  • At high-entropy branching points where probabilities diverge, it distributes probability evenly among candidate tokens, preserving exploration diversity while sharpening the distribution.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)