Purified OPSD: On-Policy Self-Distillation Without Losing How to Think
- Published
- Source
- arXiv
- Paper number
- 556
- Field
- AI / General
- arXiv ID
- 2607.02234
Key points
- OPSD fails consistently on four long-CoT models: Qwen3-8B, Qwen3-4B, R1-Distill-7B, and OLMo-7B.
- Decomposing the teacher signal shows that the reference-induced component dominates both the update direction and size, while the question-conditioned component is orthogonal or opposite.
- Epistemic markers such as 'Wait' and 'Maybe' explode or collapse during OPSD training, hurting reflective reasoning ability.
- The PMI target distribution, log πT - log πref, removes non-transferable components and is robust to β and c.
- On AIME 2024 and 2025, it consistently improves over OPSD-Standard and even outperforms the base model.
- The method remains stably improved across soft-clipping thresholds c in {5, 10, 20} and correction strengths β in {0.5, 1, 2}.
Paper links
External research summaries. These are not HDATF publications or measured product results.