LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training
- Published
- Source
- arXiv
- Paper number
- 1006
- Field
- Computer Vision
- arXiv ID
- 2608.24845
Key points
- It filtered CommonCrawl records for YouTube, Vimeo, and Dailymotion links to assemble 1.3 billion links, then downloaded 80 million of them to obtain 10 million hours of video.
- From these, it selected 2.4 million videos and split them at scene transitions into 55 million clips, automatically generating video descriptions with Qwen3-VL-2B-Instruct and audio descriptions with Audio Flamingo 3.
- After seeing 50 million training samples, BVD-V-10M and BVD-V-50M achieved overall averages 3.3 and 4.0 points higher, respectively, than the more heavily filtered InternVid-10M-FLT.
- As an open resource that handles video, audio, and images together, it is readily useful for tasks where embedding quality matters, such as retrieval and representation learning.
- However, all captions are short and automatically generated by small models, so they can be limited in expression and inherit model biases. In fact, ImageNet classification scores were lower than with existing web-image data.
Paper links
External research summaries. These are not HDATF publications or measured product results.