Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos
AuthorsSreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze, Nishit Anand, Siddharth Gururani, Hanrong Ye, Pritam Biswas, Yuanhang Su, Ehsan Hosseini-Asl, Sang-gil Lee, Zhifeng Kong, Jaehyeon Kim, Sungwon Kim, S Sakshi, Ramani Duraiswami, Dinesh Manocha, Andrew Tao, Mohammad Shoeybi, Bryan Catanzaro, Ming-Yu Liu, Wei Ping
Resources
AV-Flamingo is an open multimodal language model that learns to reason across audio, visuals, and long videos while grounding its explanations in time.
Key results
Approximate total of caption and question-answer instances in the joint audio-visual training dataset.
Question-answer pairs included in Audio-Visual-Skills for cross-modal reasoning supervision.
Maximum audio-visual duration supported during long-context training, in minutes.
AVF-Think accuracy on the long and complex audio-visual MMOU benchmark.
AVF-Instruct accuracy on the DailyOmni omni-modal understanding benchmark.
AVF-Instruct accuracy on Video-MME when subtitles are unavailable.
What the paper found
Researchers at NVIDIA and the University of Maryland introduce Nemotron-Labs-Audio-Visual Flamingo, or AV-Flamingo, a fully open audio-visual language model built for long, complex videos rather than short clips. Starting from OmniVinci and using Qwen2.5-7B as its reasoning backbone, the system combines a SigLip vision encoder, AF-Whisper for long-form audio, timestamp-based cross-modal interleaving, Constrained Rotary Time Embeddings, and streaming text-to-speech. Its main dataset, Audio-Visual-Skills, provides approximately 7M caption and question-answer training instances, including 4.8M QA pairs, covering temporal, compositional, counting, causal, and audio-visual alignment skills. A three-stage curriculum progresses from short-context perception to videos as long as 15 minutes with 32K-token contexts, followed by AV-Think post-training with Temporal Audio-Visual Interleaved Chain-of-Thought, or TAVIT, which anchors intermediate reasoning steps to timestamps and uses supervised fine-tuning plus GRPO reinforcement learning. Across more than 15 benchmarks, AVF-Instruct reaches 72.4 on DailyOmni, 60.2 on MMOU, and 70.7 on Video-MME without subtitles, outperforming similarly sized open models and challenging larger systems such as Gemini and GPT-4o, especially on long-horizon audio-visual reasoning.
Original abstract
We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over long and complex real-world (audio-visual) videos. To support this, we make three key contributions: (i) Audio-Visual-Skills, a large-scale collection of real-world videos with ~7M caption and question-answer training instances designed to emphasize temporal, compositional, and cross-modal audio-visual reasoning; (ii) a novel three-stage curriculum that progressively trains the model from short-range perception to long-horizon multi-event reasoning; and (iii) Temporal Audio-Visual Interleaved Chain-of-Thought, a reasoning framework that explicitly grounds intermediate reasoning steps to timestamps in long audio-visual streams, improving temporal alignment and interpretability. Extensive experiments across 15+ audio-visual, omni-modal, audio, and vision benchmarks show that AV-Flamingo outperforms similarly sized open models by clear margins and remains highly competitive with, and in some cases surpasses, much larger open-weight and closed models, particularly on long and complex real-world audio-visual understanding and reasoning tasks. Beyond benchmark performance, AV-Flamingo exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.