Visual Contrastive Self-Distillation
- Published
- Source
- arXiv
- Paper number
- 705
- Field
- Computer Vision
- arXiv ID
- 2607.21556
Key points
- It learns visual dependence by comparing the difference in token probabilities between the original image and an image with the content removed.
- It needs no external teacher, privileged ground truth, visual evidence signal, or reasoning trajectory.
- It adds no inference-time cost, because the EMA teacher is evaluated twice only during training.
- On Qwen3-VL, it improves by +4.77 percentage points at 2B, +1.86 points at 4B, and +3.75 points at 8B.
- It consistently outperforms previous self-distillation methods across 7 vision-language benchmarks.
Paper links
External research summaries. These are not HDATF publications or measured product results.