Latent On-Policy Self-Distillation
- Published
- Source
- arXiv
- Paper number
- 907
- Field
- Machine Learning
- arXiv ID
- 2608.13040
Key points
- It changes the design so that the training process, not a human, decides what to show the teacher, meaning the experience representation itself is optimized end to end.
- After training, only the student model remains, and performance is preserved even if the retrieval database, compressor, and latent tokens are all removed, because the experience has been internalized into the policy.
- It beats GRPO and Skill-SD with less than 30 percent of their rollout budget, which greatly reduces the compute cost of self-evolution.
- The key ingredient is a privileged margin constraint that prevents the teacher from collapsing into the student; ablations show that without this constraint the supervision signal disappears.
- The student's behavior also changes: it shifts from calling tools in a bursty, all-at-once way to executing plans step by step, with calls per step falling from 3.50 to 1.11.
Paper links
External research summaries. These are not HDATF publications or measured product results.