LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training
TL;DR - LAION-BVD is an open multimodal pre-training dataset comprising 80 million downloaded videos totaling 10 million hours, sourced from 1.3 billion CommonCrawl-discovered URLs. It substantially expands publicly accessible data for training video, audio, and image models.
- Content-aware scene detection produces clips with synthetically generated video and audio captions.
- Models trained on the data show competitive video-text and audio-text benchmark performance, improving consistently with training and model scale.
- Scene-changing video frames provide image-text data with a different visual distribution from conventional web-image corpora.
- Models trained on the extracted frames achieve strong image-text retrieval performance.