LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding

Published
Source
arXiv
Paper number
248
Field
Computer Vision
arXiv ID
2605.27365

Key points

  • Under causal attention for NTP, standard tokens can only see tokens that come before them.
  • Fast mode uses parallel decoding for everything. It offers the highest throughput, but it can sometimes fail in very dense or complex scenes.
  • Slow mode falls back to standard sequential NTP. It is more robust and precise, but much slower.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)