What Does Privileged Information Add to On-Policy Self-Distillation?

Published
Source
arXiv
Paper number
1097
Field
LLMs / NLP
arXiv ID
2609.20612

Key points

  • Built AMPLE-MATH, a dataset of 5,319 math problems each with six views that share the same answer but differ in explanation level, making it possible to isolate and measure the effect of privileged information.
  • In Qwen3-1.7B, most of the distillation gains appeared even without answer information, and the added benefit was modest, peaking with a polished solution.
  • In SmolLM3-3B, the full solution trace added about two percentage points at step 50.
  • With problems, answers, and evaluation held fixed, switching only the student's rollout mode from short direct responses to long thinking-enabled ones turned gains into losses in both model families.
  • Changing the teacher's token-level supervision signals often left student behavior largely unchanged, and 76% of the apparent gain from widening the loss window actually came from earlier stopping.
  • The paper concludes that OPSD mainly improves access to existing reasoning capabilities, and that the value of privileged information should be judged by what it adds to this transfer, not by how much of the solution it reveals.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)