NTH

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

AuthorsSreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze, Nishit Anand, Siddharth Gururani, Hanrong Ye, Pritam Biswas, Yuanhang Su, Ehsan Hosseini-Asl, Sang-gil Lee, Zhifeng Kong, Jaehyeon Kim, Sungwon Kim, S Sakshi, Ramani Duraiswami, Dinesh Manocha, Andrew Tao, Mohammad Shoeybi, Bryan Catanzaro, Ming-Yu Liu, Wei Ping

July 27, 2026 2 min read
Watch on YouTube
The one-line take

AV-Flamingo is an open multimodal language model that learns to reason across audio, visuals, and long videos while grounding its explanations in time.

Key results

7M
Audio-Visual-Skills training instances

Approximate total of caption and question-answer instances in the joint audio-visual training dataset.

4.8M
Audio-Visual-Skills QA pairs

Question-answer pairs included in Audio-Visual-Skills for cross-modal reasoning supervision.

15
Maximum video duration

Maximum audio-visual duration supported during long-context training, in minutes.

60.2
MMOU accuracy

AVF-Think accuracy on the long and complex audio-visual MMOU benchmark.

72.4
DailyOmni accuracy

AVF-Instruct accuracy on the DailyOmni omni-modal understanding benchmark.

70.7
Video-MME accuracy without subtitles

AVF-Instruct accuracy on Video-MME when subtitles are unavailable.

What the paper found

Researchers at NVIDIA and the University of Maryland introduce Nemotron-Labs-Audio-Visual Flamingo, or AV-Flamingo, a fully open audio-visual language model built for long, complex videos rather than short clips. Starting from OmniVinci and using Qwen2.5-7B as its reasoning backbone, the system combines a SigLip vision encoder, AF-Whisper for long-form audio, timestamp-based cross-modal interleaving, Constrained Rotary Time Embeddings, and streaming text-to-speech. Its main dataset, Audio-Visual-Skills, provides approximately 7M caption and question-answer training instances, including 4.8M QA pairs, covering temporal, compositional, counting, causal, and audio-visual alignment skills. A three-stage curriculum progresses from short-context perception to videos as long as 15 minutes with 32K-token contexts, followed by AV-Think post-training with Temporal Audio-Visual Interleaved Chain-of-Thought, or TAVIT, which anchors intermediate reasoning steps to timestamps and uses supervised fine-tuning plus GRPO reinforcement learning. Across more than 15 benchmarks, AVF-Instruct reaches 72.4 on DailyOmni, 60.2 on MMOU, and 70.7 on Video-MME without subtitles, outperforming similarly sized open models and challenging larger systems such as Gemini and GPT-4o, especially on long-horizon audio-visual reasoning.

Original abstract

We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over long and complex real-world (audio-visual) videos. To support this, we make three key contributions: (i) Audio-Visual-Skills, a large-scale collection of real-world videos with ~7M caption and question-answer training instances designed to emphasize temporal, compositional, and cross-modal audio-visual reasoning; (ii) a novel three-stage curriculum that progressively trains the model from short-range perception to long-horizon multi-event reasoning; and (iii) Temporal Audio-Visual Interleaved Chain-of-Thought, a reasoning framework that explicitly grounds intermediate reasoning steps to timestamps in long audio-visual streams, improving temporal alignment and interpretability. Extensive experiments across 15+ audio-visual, omni-modal, audio, and vision benchmarks show that AV-Flamingo outperforms similarly sized open models by clear margins and remains highly competitive with, and in some cases surpasses, much larger open-weight and closed models, particularly on long and complex real-world audio-visual understanding and reasoning tasks. Beyond benchmark performance, AV-Flamingo exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis