Rethinking On-Policy Distillation of Large Language Models II: One Training Example
- Published
- Source
- arXiv
- Paper number
- 1066
- Field
- AI / General
- arXiv ID
- 2609.04172
Key points
- Even with just 1 question for distillation training, performance kept improving for hundreds of steps, recovering 72% of the full-data effect on math tasks.
- It proposed 'state coverage' as an explanatory measure. 1 question achieved 71.5% coverage, while 16 semantically different questions achieved 98.9% coverage and matched full-data training performance.
- The absorption rate (the fraction of the remaining teacher-student gap closed by a single update) declined at a similar rate whether training used 1 question or 17,000. This suggests that the bottleneck is the learning algorithm rather than the amount of data.
- Training on nearly content-free templates or unrelated-domain conversation data (WildChat) had effects similar to training on real questions. The argument is that a question's only role is to 'trigger thinking.'
- Over 1000 steps using the same single question, distillation (OPD) produced at least twice the performance gain of reinforcement learning (RLVR). Outcome rewards are exhausted once the answer is correct, whereas token-level supervision continues to provide a signal.
Paper links
External research summaries. These are not HDATF publications or measured product results.