Visual Contrastive Self-Distillation

Published
Source
arXiv
Paper number
705
Field
Computer Vision
arXiv ID
2607.21556

Key points

  • It learns visual dependence by comparing the difference in token probabilities between the original image and an image with the content removed.
  • It needs no external teacher, privileged ground truth, visual evidence signal, or reasoning trajectory.
  • It adds no inference-time cost, because the EMA teacher is evaluated twice only during training.
  • On Qwen3-VL, it improves by +4.77 percentage points at 2B, +1.86 points at 4B, and +3.75 points at 8B.
  • It consistently outperforms previous self-distillation methods across 7 vision-language benchmarks.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)