OPD-V: Visual On-Policy Self-Distillation with Modality Balance
- Published
- Source
- arXiv
- Paper number
- 824
- Field
- Computer Vision
- arXiv ID
- 2608.05131
Key points
- It quantifies modality imbalance in multimodal models and proposes a new paradigm that uses this text bias as privileged information.
- It defines a modality-balanced trust region from the logit difference between a zoom-in image as the positive teacher and a masked image as the negative teacher, and it uses this region for selective distillation.
- On Qwen3.5-4B, it improves the average score from 64.3% to 80.0%, which also beats the 397B model at 77.4%.
- It shortens response length by 74.5% while improving accuracy, which reduces training time per sample by 25% to 32%.
- It shows consistent gains across six benchmarks, four backbones, and five post-processing methods.
Paper links
External research summaries. These are not HDATF publications or measured product results.