Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos
TL;DR - AV-Flamingo is a fully open, state-of-the-art audio-visual large language model built for joint understanding and reasoning over audio, images, and long-form videos, targeting the underserved regime of long and complex real-world content.
- Introduces Audio-Visual-Skills, a large-scale dataset of real-world videos with ~7M caption and QA training instances emphasizing temporal, compositional, and cross-modal reasoning.
- Uses a three-stage curriculum that progresses from short-range perception to long-horizon multi-event reasoning.
- Proposes Temporal Audio-Visual Interleaved Chain-of-Thought, grounding intermediate reasoning steps to timestamps for better temporal alignment and interpretability.
- Reports outperforming similarly sized open models across 15+ benchmarks and remaining competitive with (sometimes surpassing) much larger open-weight and closed models, especially on long/complex tasks.