Latent On-Policy Self-Distillation

Published
Source
arXiv
Paper number
907
Field
Machine Learning
arXiv ID
2608.13040

Key points

  • It changes the design so that the training process, not a human, decides what to show the teacher, meaning the experience representation itself is optimized end to end.
  • After training, only the student model remains, and performance is preserved even if the retrieval database, compressor, and latent tokens are all removed, because the experience has been internalized into the policy.
  • It beats GRPO and Skill-SD with less than 30 percent of their rollout budget, which greatly reduces the compute cost of self-evolution.
  • The key ingredient is a privileged margin constraint that prevents the teacher from collapsing into the student; ablations show that without this constraint the supervision signal disappears.
  • The student's behavior also changes: it shifts from calling tools in a bursty, all-at-once way to executing plans step by step, with calls per step falling from 3.50 to 1.11.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)