Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation

Published
Source
arXiv
Paper number
198
Field
Computer Vision
arXiv ID
2605.18740

Key points

  • Question generation: For each small region, the MLLM generates a question q that can be answered only by seeing the details inside that region.
  • Grounding to the full image: The region's bounding box is overlaid on the full image I to form the student input x. A spatial instruction, such as 'look at the red box', is added to the question so the student knows where to focus.
  • Crop generation: The region is cropped and resized, often by 2x, to form the teacher input x'.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)