AsyncOPD: How Stale Can On-Policy Distillation Be?

Published
Source
arXiv
Paper number
496
Field
Machine Learning
arXiv ID
2606.24143

Key points

  • AsyncOPD systematically studies the stale-data problem in asynchronous On-Policy Distillation, or OPD.
  • It shows that the KL direction changes the stale-data problem, where teacher-weighted forward KL is robust and student-weighted reverse KL is fragile, and it shows that recomputing advantages at training time is the most effective fix for reverse-KL OPD.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)