What Does Privileged Information Add to On-Policy Self-Distillation?
- Published
- Source
- arXiv
- Paper number
- 1097
- Field
- LLMs / NLP
- arXiv ID
- 2609.20612
Key points
- Built AMPLE-MATH, a dataset of 5,319 math problems each with six views that share the same answer but differ in explanation level, making it possible to isolate and measure the effect of privileged information.
- In Qwen3-1.7B, most of the distillation gains appeared even without answer information, and the added benefit was modest, peaking with a polished solution.
- In SmolLM3-3B, the full solution trace added about two percentage points at step 50.
- With problems, answers, and evaluation held fixed, switching only the student's rollout mode from short direct responses to long thinking-enabled ones turned gains into losses in both model families.
- Changing the teacher's token-level supervision signals often left student behavior largely unchanged, and 76% of the apparent gain from widening the loss window actually came from earlier stopping.
- The paper concludes that OPSD mainly improves access to existing reasoning capabilities, and that the value of privileged information should be judged by what it adds to this transfer, not by how much of the solution it reveals.
Paper links
External research summaries. These are not HDATF publications or measured product results.