Pass the Baton: Trajectory-Relayed On-Policy Distillation

Published
Source
arXiv
Paper number
750
Field
LLMs / NLP
arXiv ID
2607.26057

Key points

  • It discovers a teacher-student asymmetry in direction switching: the teacher reflects and changes direction, but the student just keeps going straight.
  • It experimentally confirms that even 0.35% teacher tokens can raise accuracy by +7.23 points.
  • It limits the number and length of teacher interventions through a relay budget so the student policy does not drift too far away.
  • It cuts the average training trajectory length by more than 50%, from 4,658 to 2,296 tokens, while still improving accuracy.
  • A single regression objective makes stable training possible without VSD or GAN losses.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)