ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

Published
Source
arXiv
Paper number
778
Field
Computer Vision
arXiv ID
2607.28627

Key points

  • It finds that the attention, or query-key, scores of a VLM are poorly suited for visual information retrieval, with an average recall@1 of only 5.1 percent.
  • It adds a single learnable token, ReToken, to explicitly retrieve relevant frames in value space.
  • On Visual Haystacks, it improves Qwen3VL-8B by 13.4 points and InternVL3.5 by 12.4 points, which is more than a 20 percent relative gain.
  • Even when trained only on image QA, it transfers zero-shot to long videos and improves LVBench by 8.0 points.
  • It is lightweight enough to train and run on a single H100, and the added retrieval latency is about 0.4 seconds.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)