VideoRAG: Retrieval-Augmented Generation over Video Corpus
- Published
- Source
- arXiv
- Paper number
- 018
- Field
- RAG / Video
- arXiv ID
- 2501.05874
Key points
- Large language models and large vision-language models often generate inaccurate or stale information because they rely on parametric knowledge.
- Conventional retrieval-augmented generation, or RAG, focuses mainly on text data and includes static images only in a limited way, while video is often converted to text and loses rich multimodal information such as temporal dynamics and visual context.
- Long videos challenge LVLMs because of computational infeasibility and limited context windows, and many videos do not have associated text such as subtitles.
- VideoRAG embeds both the query and the video, using frames and subtitles, with an LVLM and selects the top-k relevant videos by cosine similarity to enable dynamic video retrieval from a large corpus.
- For response generation, the framework connects the selected video frames and associated text, either subtitles or generated auxiliary text, with the user query and feeds this comprehensive multimodal input into the LVLM.
- An adaptive frame selection mechanism with k-means++ clustering identifies the most useful and least redundant video frames, and when subtitles are unavailable it uses ASR to generate auxiliary text from the audio track.
Paper links
External research summaries. These are not HDATF publications or measured product results.