VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
AuthorsXinhao Li, Yuhan Zhu, Xiangyu Zeng, Yuhao Dong, Haoning Wu, Zhiqiu Zhang, Yuandong Yang, Changlian Ma, Qingyu Zhang, Yansong Shi, Xinyu Chen, Haoran Chen, Zizheng Huang, Jun Zhang, Kun Ouyang, Lin Sui, Ziang Yan, Yicheng Xu, Chenting Wang, Yinan He, Hongjie Zhang, Yi Wang, Yu Qiao, Yali Wang, Ziwei Liu, Kai Chen, Limin Wang
Resources
VideoChat3 is a fully open, efficient 4B-parameter video-language model designed to understand everything from short clips to long and streaming videos.
Key results
VideoChat3 parameter count
Spatiotemporal visual-token compression using four-frame temporal pooling and 2×2 spatial merging
Combined multimodal instruction samples across VideoChat3-Academic2M, VideoChat3-LV116K, and VideoChat3-OL617K
Average F1 for proactive streaming response timing
VideoChat3 total inference latency, compared with 44.449s for Qwen3-VL
What the paper found
VideoChat3, developed by researchers at Nanjing University, Shanghai AI Laboratory, Nanyang Technological University, and Peking University, is a fully open 4B-parameter video multimodal large language model designed to unify short-video perception, long-context reasoning, temporal grounding, and live streaming interaction. Its central architectural contribution is the Inflated 3D Vision Transformer, or I3D-ViT, which extends image-based self-attention into chunk-wise spatiotemporal attention and performs temporal pooling before visual tokens reach the language model; with four-frame chunks and 2×2 spatial merging, it achieves 16× spatiotemporal compression. Adaptive Frame Resolution adds a state-controlled streaming policy: Silence uses a 224² pixel quota, Standby raises the next window to 448², and Response resumes low-cost monitoring. The training pipeline produces roughly 3M multimodal instruction samples through VideoChat3-Academic2M, VideoChat3-LV116K, and VideoChat3-OL617K, using evidence-grounded annotation enhancement, long-video segmentation, clue verification, and a balanced state-transition loss. VideoChat3 scores 61.7 on MotionBench, 75.6 on TempCompass, and 35.5 average F1 on OVO-Timing, outperforming similarly sized open models such as Qwen3-VL-4B and Molmo2-4B while remaining competitive with proprietary systems including OpenAI’s GPT-5 and Google’s Gemini. Efficiency improves sharply with longer inputs: on 2048 frames, total latency falls from 44.449s for Qwen3-VL to 20.412s for VideoChat3. The authors release weights, code, training recipes, datasets, and data-construction pipelines, making the system unusually reproducible.
Original abstract
Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current open-source models remain limited in several ways. They often struggle to generalize across diverse video types, making them effective only in specific domains. High computational demands further restrict their efficiency and scalability. Moreover, most models are only partially open, with key components such as training code, strategy, or datasets unavailable, which hinders reproducibility and slows community-driven development. To address these issues, we introduce VideoChat3, a fully open, efficient, and generalist video-centric MLLM. VideoChat3 advances video understanding through two complementary designs. For efficiency, we introduce Inflated 3D Vision Transformer (I3D-ViT) and Adaptive Frame Resolution for Streaming Video Perception, which enables efficient spatiotemporal representation and reduces the cost of processing video inputs during training and inference. For effectiveness, we develop a scalable video data synthesis pipeline that curates three diverse, high-quality training datasets: VideoChat3-Academic2M, VideoChat3-LV116K, and VideoChat3-OL617K, covering general, long-form, and streaming video scenarios, improving the model's generalization across domains. By integrating these designs, VideoChat3 achieves a rare balance of broad generalization and computational efficiency. Experiments across general, long-form, and streaming benchmarks demonstrate that VideoChat3 surpasses prior open-source models with equal or larger parameter counts with only 4B parameters and higher efficiency.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.