Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
- Published
- Source
- arXiv
- Paper number
- 198
- Field
- Computer Vision
- arXiv ID
- 2605.18740
Key points
- Question generation: For each small region, the MLLM generates a question q that can be answered only by seeing the details inside that region.
- Grounding to the full image: The region's bounding box is overlaid on the full image I to form the student input x. A spatial instruction, such as 'look at the red box', is added to the question so the student knows where to focus.
- Crop generation: The region is cropped and resized, often by 2x, to form the teacher input x'.
Paper links
External research summaries. These are not HDATF publications or measured product results.