Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models
- Published
- Source
- arXiv
- Paper number
- 939
- Field
- LLMs / NLP
- arXiv ID
- 2608.16647
Key points
- On-policy distillation reduces the difference between the teacher's and student's next-token distributions at each point in an answer generated by the student. The student therefore receives guidance in the contexts it actually visits during training.
- For BigMath, the teacher selected 25,000 problems from each of three groups: problems answered correctly on all four attempts, problems answered incorrectly on all four attempts, and randomly selected problems. Across several teacher and student combinations, the three groups produced nearly identical final mathematics accuracy.
- A teacher derived from the same base model brought the student close to the teacher's level in Chinese, long-horizon mathematics, coding, and science, even when training covered only English and short-horizon mathematics. A stronger teacher with a different base model produced much weaker transfer across those areas.
- Changing the ratio of data from two teachers caused the student's scores across several areas to move together toward the original capability profile of the teacher used more heavily. For the first teacher pair, the average score on BeyondAIME and OlymMATH ranged from 25.1% to 27.1% across mixture conditions.
- The experiments were limited to mathematics, coding, science, and instruction-following models. The multi-teacher study also covered only two teachers with fixed domain-based routing. The authors warn that domain partitioning should not be treated as a safety boundary because undesirable teacher behavior may also spread across domains.
Paper links
External research summaries. These are not HDATF publications or measured product results.