On-Policy Delta Distillation
- Published
- Source
- arXiv
- Paper number
- 651
- Field
- Machine Learning
- arXiv ID
- 2607.15161
Key points
- It argues that using the difference between a teacher's pretraining and post-training states, rather than the teacher's final output, transfers reasoning ability more accurately.
- It consistently outperforms existing on-policy distillation methods across 14 benchmarks in math, science, and code.
- It is validated broadly on recent models such as Qwen3, from 1.7B to 8B, and Gemma4, which makes the result credible.
- The additional compute cost is only 8% to 28%, so the method can be added to existing distillation pipelines easily.
- Its analysis shows that the delta signal concentrates on reasoning tokens such as hence and however, while suppressing exploratory tokens such as try and verify.
Paper links
External research summaries. These are not HDATF publications or measured product results.