AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning

Published
Source
arXiv
Paper number
586
Field
Computer Vision
arXiv ID
2607.07033

Key points

  • Its sequential design first anchors relevance and then expands context, which prioritizes non-redundant query evidence.
  • It automatically chooses the anchor size from each token's novelty profile rather than using a fixed ratio.
  • The architecture-aware design applies to both CLIP-aligned and unaligned models, as well as image and video VLMs.
  • On LLaVA-NeXT-7B, it retains 97.6 percent of full performance with only 160 of 2,880 tokens, or 5.6 percent.
  • At extreme compression with 32 tokens, it outperforms prior relevance-diversity integrated methods by 3.7 to 5.7 percentage points.
  • It is a lightweight framework that requires no training and can be applied directly to existing VLMs without architectural changes.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)