Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos
- Published
- Source
- arXiv
- Paper number
- 659
- Field
- Computer Vision
- arXiv ID
- 2607.16107
Key points
- It releases the AV-Skills dataset, consisting of about 7 million caption and QA instances.
- It designs a three-stage curriculum that progresses from short-video recognition to reasoning over long videos of up to 15 minutes.
- TAVIT precisely maps intermediate reasoning steps to video and audio timestamps.
- It achieves 60.2% on MMOU, the best performance among comparable open-source models.
Paper links
External research summaries. These are not HDATF publications or measured product results.