LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
- Published
- Source
- arXiv
- Paper number
- 248
- Field
- Computer Vision
- arXiv ID
- 2605.27365
Key points
- Under causal attention for NTP, standard tokens can only see tokens that come before them.
- Fast mode uses parallel decoding for everything. It offers the highest throughput, but it can sometimes fail in very dense or complex scenes.
- Slow mode falls back to standard sequential NTP. It is more robust and precise, but much slower.
Paper links
External research summaries. These are not HDATF publications or measured product results.