On-Policy Self-Distillation without Any Supervision

Published
Source
arXiv
Paper number
842
Field
Machine Learning
arXiv ID
2608.06296

Key points

  • It achieves self-distillation using only the model's own generations, without external answers, environment feedback, or a larger model.
  • It forms a pseudo-answer by majority vote and uses the shortest agreeing rollout as the teacher.
  • It applies distillation to the prefix of incorrect rollouts so that the model corrects the wrong part precisely.
  • On math benchmarks, it improves by 8.5 to 10.7 points over the base model without reasoning mode and by 3.2 to 2.3 points over OPSD.
  • Forward KL works best, reverse KL diverges during training, and JSD is ineffective.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)