🛰️ Daily AI Frontier
‹ back to 2026-08-26

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

Research Multimodal & Generative

Ranking

Overall 86
Content 95
Popularity 65

Observed public metrics from 1 member.

Merged summary

TL;DR - LAION-BVD is an open multimodal pre-training dataset comprising 80 million downloaded videos totaling 10 million hours, sourced from 1.3 billion CommonCrawl-discovered URLs. It substantially expands publicly accessible data for training video, audio, and image models.

  • Content-aware scene detection produces clips with synthetically generated video and audio captions.
  • Models trained on the data show competitive video-text and audio-text benchmark performance, improving consistently with training and model scale.
  • Scene-changing video frames provide image-text data with a different visual distribution from conventional web-image corpora.
  • Models trained on the extracted frames achieve strong image-text retrieval performance.

Sources (1)

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

arXiv cs.CV Andreas Hochlehnert, Marianna Nezhurina, Mehdi Cherti, Andrej Radonjic, Thaddäus Wiedemer, Christoph Schuhmann, Romain Beaumont, Wieland Brendel, Bernhard Schölkopf, A. Sophia Koepke, Jenia Jitsev, Matthias Bethge 2026-08-25 arXiv:2608.24845
Public signals Hugging Face upvotes 16
Providers: Hugging Face · Upvotes 16 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:28:40.640675 UTC

TL;DR - LAION-BVD is an open multimodal pre-training dataset comprising 80 million downloaded videos totaling 10 million hours, sourced from 1.3 billion CommonCrawl-discovered URLs. It substantially expands publicly accessible data for training video, audio, and image models.

  • Content-aware scene detection produces clips with synthetically generated video and audio captions.
  • Models trained on the data show competitive video-text and audio-text benchmark performance, improving consistently with training and model scale.
  • Scene-changing video frames provide image-text data with a different visual distribution from conventional web-image corpora.
  • Models trained on the extracted frames achieve strong image-text retrieval performance.
item →