AsyncOPD: How Stale Can On-Policy Distillation Be?
- Published
- Source
- arXiv
- Paper number
- 496
- Field
- Machine Learning
- arXiv ID
- 2606.24143
Key points
- AsyncOPD systematically studies the stale-data problem in asynchronous On-Policy Distillation, or OPD.
- It shows that the KL direction changes the stale-data problem, where teacher-weighted forward KL is robust and student-weighted reverse KL is fragile, and it shows that recomputing advantages at training time is the most effective fix for reverse-KL OPD.
Paper links
External research summaries. These are not HDATF publications or measured product results.