LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training
cs.CV, cs.AI, cs.LG
Submitted: 2026-08-25
Updated: 2026-08-25
Code: https://github.com/rom1504/cc2dataset
License: http://creativecommons.org/licenses/by/4.0/
The gist: We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl.
Terminology
Abstract
We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a total duration of 10 million hours. The dataset is designed for multimodal pre-training across the video, audio, and image modalities. Using content-aware scene detection, we extract clips for which we synthetically generate video and audio captions. Models trained on these data achieve competitive performance on standard video-text and audio-text benchmarks, with consistent improvements as training or model scale increases. Additionally, we explore video frames as an alternative source of image-text data by extracting scene-changing frames. These frames exhibit a visual distribution distinct from standard web image corpora, and models trained on this dataset achieve strong image-text retrieval performance. We release LAION-BVD to the research community. It significantly expands open access to multimodal videos at an unprecedented scale.
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Qwen3-VL Technical Report
- Datasheet for the Pile
- MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
- PaLI: A Jointly-Scaled Multilingual Language-Image Model
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- RedCaps: web-curated image-text data created by the people, for the people
- Scalable Vision Language Model Training via High Quality Data Curation
- mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos
- Exploring the Frontier of Vision-Language Models: A Survey of Current Methodologies and Future Directions
- Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
- Non-Autoregressive Neural Machine Translation
- Non-Autoregressive Image Captioning with Counterfactuals-Critical Multi-Agent Learning
- WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models
- CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models