Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation

Published
Source
arXiv
Paper number
952
Field
Machine Learning
arXiv ID
2608.19098

Key points

  • Even oracle routing based on ground-truth domain labels, which assigns each prompt to the correct teacher, fails to integrate the capabilities successfully. The bottleneck is capability integration itself, not routing.
  • Teacher disagreement was not the main cause. The problem was a distorted token budget: instruction-following tasks made up 20.3% of input prompts but received only 0.99% of gradient tokens.
  • The study separates the distortion into three causes: differences in response length across domains, different convergence rates across domains, and stale student reward signals caused by rollout reuse.
  • Combining token-share balancing, gap-aware dynamic budget allocation, and student reward refresh raised the improvement recovery rate from 35.6% to 83.4%.
  • The authors released the full training recipe, training trajectories, and evaluation suite, all reproducible on an academic setup with eight A100 80GB GPUs.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)