Purified OPSD: On-Policy Self-Distillation Without Losing How to Think

Published
Source
arXiv
Paper number
556
Field
AI / General
arXiv ID
2607.02234

Key points

  • OPSD fails consistently on four long-CoT models: Qwen3-8B, Qwen3-4B, R1-Distill-7B, and OLMo-7B.
  • Decomposing the teacher signal shows that the reference-induced component dominates both the update direction and size, while the question-conditioned component is orthogonal or opposite.
  • Epistemic markers such as 'Wait' and 'Maybe' explode or collapse during OPSD training, hurting reflective reasoning ability.
  • The PMI target distribution, log πT - log πref, removes non-transferable components and is robust to β and c.
  • On AIME 2024 and 2025, it consistently improves over OPSD-Standard and even outperforms the base model.
  • The method remains stably improved across soft-clipping thresholds c in {5, 10, 20} and correction strengths β in {0.5, 1, 2}.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)