Describe Anything: Detailed Localized Image and Video Captioning

Published
Source
arXiv
Paper number
063
Field
Vision-Language
arXiv ID
2504.16072

Key points

  • Existing vision-language models struggle to generate detailed and accurate captions for specific regions in images and videos.
  • High-quality datasets for training detailed local captioning models are limited.
  • Current evaluation metrics often rely on reference captions that lack comprehensive detail.
  • It developed the DAM architecture for multi-granularity region captioning, with focal prompts and a localized vision backbone.
  • It built the SSL-based data pipeline DLC-SDP using segmentation datasets and unlabeled web images.
  • It introduced DLC-Bench for evaluating detailed local captions without reference captions.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)