Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation
- Published
- Source
- arXiv
- Paper number
- 870
- Field
- Machine Learning
- arXiv ID
- 2608.09228
Key points
- It designs a controlled experiment, OP2SD, to distinguish whether OPSD's gains come from answer transfer or context-induced behavior.
- Even when the teacher is given solutions to other math problems instead of the correct answer, the improvement is similar to OPSD.
- When the context comes from physics rather than math, the effect weakens, showing that domain alignment in the context is necessary.
- It confirms the same pattern across three Qwen3 models, 1.7B, 4B, and 8B, and on the AIME and HMMT benchmarks.
- It shows that OPSD should not be interpreted only as privileged-information transfer, because context-induced behavior is also an important factor.
Paper links
External research summaries. These are not HDATF publications or measured product results.