Self-Supervised Visual On-Policy Distillation
- Published
- Source
- arXiv
- Paper number
- 916
- Field
- Computer Vision
- arXiv ID
- 2608.14144
Key points
- The teacher model sees clean images, while the student sees corrupted versions of the same images and learns to match the teacher's output distribution.
- The best setting lowers the student's image resolution to 0.3x to 0.6x of the original and adds Gaussian noise with 50% probability.
- Training uses Qwen3.5 4B and 9B, and evaluation covers six visual-recognition tasks and three mathematical-reasoning tasks.
- On 12,000 FineVision questions, the 4B model's average across six visual-recognition tasks rises from 70.68% to 77.44%, a gain of 6.76 points.
- Using the same image for teacher and student reaches only 70.52% on average, while the full method with information differences reaches 76.35%. Some analyses use shorter output limits than the main evaluation.
Paper links
External research summaries. These are not HDATF publications or measured product results.