Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

Published
Source
arXiv
Paper number
659
Field
Computer Vision
arXiv ID
2607.16107

Key points

  • It releases the AV-Skills dataset, consisting of about 7 million caption and QA instances.
  • It designs a three-stage curriculum that progresses from short-video recognition to reasoning over long videos of up to 15 minutes.
  • TAVIT precisely maps intermediate reasoning steps to video and audio timestamps.
  • It achieves 60.2% on MMOU, the best performance among comparable open-source models.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)