OPD-V: Visual On-Policy Self-Distillation with Modality Balance

Published
Source
arXiv
Paper number
824
Field
Computer Vision
arXiv ID
2608.05131

Key points

  • It quantifies modality imbalance in multimodal models and proposes a new paradigm that uses this text bias as privileged information.
  • It defines a modality-balanced trust region from the logit difference between a zoom-in image as the positive teacher and a masked image as the negative teacher, and it uses this region for selective distillation.
  • On Qwen3.5-4B, it improves the average score from 64.3% to 80.0%, which also beats the 397B model at 77.4%.
  • It shortens response length by 74.5% while improving accuracy, which reduces training time per sample by 25% to 32%.
  • It shows consistent gains across six benchmarks, four backbones, and five post-processing methods.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)