Describe Anything: Detailed Localized Image and Video Captioning
- Published
- Source
- arXiv
- Paper number
- 063
- Field
- Vision-Language
- arXiv ID
- 2504.16072
Key points
- Existing vision-language models struggle to generate detailed and accurate captions for specific regions in images and videos.
- High-quality datasets for training detailed local captioning models are limited.
- Current evaluation metrics often rely on reference captions that lack comprehensive detail.
- It developed the DAM architecture for multi-granularity region captioning, with focal prompts and a localized vision backbone.
- It built the SSL-based data pipeline DLC-SDP using segmentation datasets and unlabeled web images.
- It introduced DLC-Bench for evaluating detailed local captions without reference captions.
Paper links
External research summaries. These are not HDATF publications or measured product results.