On-Policy Self-Distillation without Any Supervision
- Published
- Source
- arXiv
- Paper number
- 842
- Field
- Machine Learning
- arXiv ID
- 2608.06296
Key points
- It achieves self-distillation using only the model's own generations, without external answers, environment feedback, or a larger model.
- It forms a pseudo-answer by majority vote and uses the shortest agreeing rollout as the teacher.
- It applies distillation to the prefix of incorrect rollouts so that the model corrects the wrong part precisely.
- On math benchmarks, it improves by 8.5 to 10.7 points over the base model without reasoning mode and by 3.2 to 2.3 points over OPSD.
- Forward KL works best, reverse KL diverges during training, and JSD is ineffective.
Paper links
External research summaries. These are not HDATF publications or measured product results.