Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation
- Published
- Source
- arXiv
- Paper number
- 952
- Field
- Machine Learning
- arXiv ID
- 2608.19098
Key points
- Even oracle routing based on ground-truth domain labels, which assigns each prompt to the correct teacher, fails to integrate the capabilities successfully. The bottleneck is capability integration itself, not routing.
- Teacher disagreement was not the main cause. The problem was a distorted token budget: instruction-following tasks made up 20.3% of input prompts but received only 0.99% of gradient tokens.
- The study separates the distortion into three causes: differences in response length across domains, different convergence rates across domains, and stale student reward signals caused by rollout reuse.
- Combining token-share balancing, gap-aware dynamic budget allocation, and student reward refresh raised the improvement recovery rate from 35.6% to 83.4%.
- The authors released the full training recipe, training trajectories, and evaluation suite, all reproducible on an academic setup with eight A100 80GB GPUs.
Paper links
External research summaries. These are not HDATF publications or measured product results.