Self-Supervised Visual On-Policy Distillation

Published
Source
arXiv
Paper number
916
Field
Computer Vision
arXiv ID
2608.14144

Key points

  • The teacher model sees clean images, while the student sees corrupted versions of the same images and learns to match the teacher's output distribution.
  • The best setting lowers the student's image resolution to 0.3x to 0.6x of the original and adds Gaussian noise with 50% probability.
  • Training uses Qwen3.5 4B and 9B, and evaluation covers six visual-recognition tasks and three mathematical-reasoning tasks.
  • On 12,000 FineVision questions, the 4B model's average across six visual-recognition tasks rises from 70.68% to 77.44%, a gain of 6.76 points.
  • Using the same image for teacher and student reaches only 70.52% on average, while the full method with information differences reaches 76.35%. Some analyses use shorter output limits than the main evaluation.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)