Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation

Published
Source
arXiv
Paper number
870
Field
Machine Learning
arXiv ID
2608.09228

Key points

  • It designs a controlled experiment, OP2SD, to distinguish whether OPSD's gains come from answer transfer or context-induced behavior.
  • Even when the teacher is given solutions to other math problems instead of the correct answer, the improvement is similar to OPSD.
  • When the context comes from physics rather than math, the effect weakens, showing that domain alignment in the context is necessary.
  • It confirms the same pattern across three Qwen3 models, 1.7B, 4B, and 8B, and on the AIME and HMMT benchmarks.
  • It shows that OPSD should not be interpreted only as privileged-information transfer, because context-induced behavior is also an important factor.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)