MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training
- Published
- Source
- arXiv
- Paper number
- 528
- Field
- LLMs / NLP
- arXiv ID
- 2606.30406
Key points
- It uses a three-stage pipeline: General SFT, Domain-specialised RL, and Multi-teacher On-Policy Distillation.
- It aligns teacher and student on student rollouts with per-token reverse KL, which gives dense supervision and remains on-policy.
- On Qwen3-30B-A3B, it improves by 5.5 points over the next best baseline and closes 91% to 95% of the teacher-student gap.
- Domain teacher training can run in parallel, which increases development throughput and decouples the domains.
- Teachers from the same origin are key to stable optimization, while using external teachers risks divergence.
- It is deployed in the MiMo-V2-Flash industrial model, which validates its utility at frontier scale.
Paper links
External research summaries. These are not HDATF publications or measured product results.