MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

Published
Source
arXiv
Paper number
528
Field
LLMs / NLP
arXiv ID
2606.30406

Key points

  • It uses a three-stage pipeline: General SFT, Domain-specialised RL, and Multi-teacher On-Policy Distillation.
  • It aligns teacher and student on student rollouts with per-token reverse KL, which gives dense supervision and remains on-policy.
  • On Qwen3-30B-A3B, it improves by 5.5 points over the next best baseline and closes 91% to 95% of the teacher-student gap.
  • Domain teacher training can run in parallel, which increases development throughput and decouples the domains.
  • Teachers from the same origin are key to stable optimization, while using external teachers risks divergence.
  • It is deployed in the MiMo-V2-Flash industrial model, which validates its utility at frontier scale.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)