Mismatch Matters: On-Policy Distillation Beyond Token Agreement

Published
Source
arXiv
Paper number
867
Field
AI / General
arXiv ID
2608.09836

Key points

  • We discovered a degenerate agreement phenomenon in which the student cheats by repeating tokens to achieve high token-level agreement with the teacher.
  • We distinguish between surplus tokens, which the teacher scores near zero, and missing tokens, which the teacher wants but the student does not produce.
  • We propose TIDE, which corrects surplus tokens with Hellinger-distance-based bounded suppression and missing tokens with analytical top-K injection.
  • It raises Avg@8 from 6.9% to 20.3% under strong teacher-student disagreement.
  • It also cuts average response length by 3.6x while greatly reducing format failures.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)