DOPD: Dual On-policy Distillation

Published
Source
arXiv
Paper number
526
Field
AI / General
arXiv ID
2606.30626

Key points

  • It defines the privilege illusion as a failure pattern in which information asymmetry from privileged information is mistaken for ability gaps.
  • Advantage-aware routing dynamically selects the supervision source, teacher or student, for each token.
  • It improves average performance by +7.5 points over Vanilla OPD on LLMs and +6.0 points on VLMs.
  • It applies strong-teacher distillation over the full vocabulary and weak-student self-supervision by token category.
  • It shows consistent gains in continual learning, OOD generalization, and training stability.
  • Token-level analysis shows that high-teacher, low-student tokens are responsible for transferring core abilities.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)