LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

Published
Source
arXiv
Paper number
1006
Field
Computer Vision
arXiv ID
2608.24845

Key points

  • It filtered CommonCrawl records for YouTube, Vimeo, and Dailymotion links to assemble 1.3 billion links, then downloaded 80 million of them to obtain 10 million hours of video.
  • From these, it selected 2.4 million videos and split them at scene transitions into 55 million clips, automatically generating video descriptions with Qwen3-VL-2B-Instruct and audio descriptions with Audio Flamingo 3.
  • After seeing 50 million training samples, BVD-V-10M and BVD-V-50M achieved overall averages 3.3 and 4.0 points higher, respectively, than the more heavily filtered InternVid-10M-FLT.
  • As an open resource that handles video, audio, and images together, it is readily useful for tasks where embedding quality matters, such as retrieval and representation learning.
  • However, all captions are short and automatically generated by small models, so they can be limited in expression and inherit model biases. In fact, ImageNet classification scores were lower than with existing web-image data.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)