DOPD: Dual On-policy Distillation
- Published
- Source
- arXiv
- Paper number
- 526
- Field
- AI / General
- arXiv ID
- 2606.30626
Key points
- It defines the privilege illusion as a failure pattern in which information asymmetry from privileged information is mistaken for ability gaps.
- Advantage-aware routing dynamically selects the supervision source, teacher or student, for each token.
- It improves average performance by +7.5 points over Vanilla OPD on LLMs and +6.0 points on VLMs.
- It applies strong-teacher distillation over the full vocabulary and weak-student self-supervision by token category.
- It shows consistent gains in continual learning, OOD generalization, and training stability.
- Token-level analysis shows that high-teacher, low-student tokens are responsible for transferring core abilities.
Paper links
External research summaries. These are not HDATF publications or measured product results.