ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
- Published
- Source
- arXiv
- Paper number
- 778
- Field
- Computer Vision
- arXiv ID
- 2607.28627
Key points
- It finds that the attention, or query-key, scores of a VLM are poorly suited for visual information retrieval, with an average recall@1 of only 5.1 percent.
- It adds a single learnable token, ReToken, to explicitly retrieve relevant frames in value space.
- On Visual Haystacks, it improves Qwen3VL-8B by 13.4 points and InternVL3.5 by 12.4 points, which is more than a 20 percent relative gain.
- Even when trained only on image QA, it transfers zero-shot to long videos and improves LVBench by 8.0 points.
- It is lightweight enough to train and run on a single H100, and the added retrieval latency is about 0.4 seconds.
Paper links
External research summaries. These are not HDATF publications or measured product results.