On-Policy Delta Distillation

Published
Source
arXiv
Paper number
651
Field
Machine Learning
arXiv ID
2607.15161

Key points

  • It argues that using the difference between a teacher's pretraining and post-training states, rather than the teacher's final output, transfers reasoning ability more accurately.
  • It consistently outperforms existing on-policy distillation methods across 14 benchmarks in math, science, and code.
  • It is validated broadly on recent models such as Qwen3, from 1.7B to 8B, and Gemma4, which makes the result credible.
  • The additional compute cost is only 8% to 28%, so the method can be added to existing distillation pipelines easily.
  • Its analysis shows that the delta signal concentrates on reasoning tokens such as hence and however, while suppressing exploratory tokens such as try and verify.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)