Pass the Baton: Trajectory-Relayed On-Policy Distillation
- Published
- Source
- arXiv
- Paper number
- 750
- Field
- LLMs / NLP
- arXiv ID
- 2607.26057
Key points
- It discovers a teacher-student asymmetry in direction switching: the teacher reflects and changes direction, but the student just keeps going straight.
- It experimentally confirms that even 0.35% teacher tokens can raise accuracy by +7.23 points.
- It limits the number and length of teacher interventions through a relay budget so the student policy does not drift too far away.
- It cuts the average training trajectory length by more than 50%, from 4,658 to 2,296 tokens, while still improving accuracy.
- A single regression objective makes stable training possible without VSD or GAN losses.
Paper links
External research summaries. These are not HDATF publications or measured product results.