Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots

Published
Source
arXiv
Paper number
869
Field
Computer Vision
arXiv ID
2608.09931

Key points

  • It proposes a three-stage verification criterion that finds visual blind spots using only the model's own counterfactual responses, without external models, annotations, or tools.
  • It performs contrastive distillation with an expanded view of the blind spot as the positive teacher and a removed view as the negative teacher.
  • On Qwen3-VL-8B, it improves all 12 benchmarks with no regression.
  • It records gains of +3.60 on OCRBench, +3.38 on fine-grained perception in MMStar, and +3.08 on logical reasoning.
  • It outperforms all six self-evolution baselines that depend on external GPT-4o supervision.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)