Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
AuthorsSenqiao Yang, Kaichen Zhang, Zhaoyang Jia, Jinghao Guo, Yifei Shen, Xinjie Zhang, Xiaoyi Zhang, Haoqing Wang, Xiao Li, Peng Zhang, Xiang An, Yin Xie, Zhening Liu, Xun Guo, Jiahao Li, Shicheng Zheng, Jinglu Wang, Zongyu Guo, Wenxuan Xie, Zihan Zheng, Yuxuan Luo, Bin Li, Yan Lu
Resources
Mage-VL turns video codecs and motion-aware tokenization into a faster multimodal model designed for real-time visual understanding.
Key results
Codec-native patch selection removes redundant visual tokens from streaming video.
Mage-ViT is trained from scratch on unlabeled images.
Mage-ViT is jointly trained on unlabeled video frames.
Mage-VL-4B uses the Qwen3-4B-Instruct language backbone.
Codec-native processing achieves up to 3.5× wall-clock speedup over uniform frame sampling.
Mage-VL-4B achieves the reported overall streaming score on OVO-Bench.
What the paper found
Microsoft’s Mage-VL addresses the streaming weakness of conventional vision-language models by making video compression part of the visual tokenizer. Its Mage-ViT encoder divides frames into 16×16 patches, densely processes anchor I-frames, and selectively retains motion- and entropy-rich patches from predicted P-frames using HEVC/H.265 motion vectors and residual energy; the same strategy also works with the neural codec DCVC-RT. Trained from scratch on 560M unlabeled images and 100M unlabeled video frames, Mage-ViT rivals encoders such as SigLIP2 that rely on far larger image-text corpora. The resulting Mage-VL-4B combines this encoder with the Qwen3-4B-Instruct language backbone and a five-stage curriculum covering captions, instruction tuning, long-video adaptation, codec-native training, and streaming alignment. A lightweight System 1 cognition gate decides when an event merits a response, while a causal System 2 decoder generates commentary only after activation. Mage-VL reduces visual token consumption by 75% and delivers up to 3.5× wall-clock inference speedup over uniform frame sampling, while matching Qwen3-VL-4B on static tasks and improving temporal grounding, video understanding, and spatial reasoning. On OVO-Bench, it reaches 64.00%, and its AI4AI pipeline—using GPT-5 rubric scoring and GitHub Copilot prompt-code refinement—improves caption quality and downstream benchmarks. The paper also reports that Zero-Vision SFT followed by multimodal reinforcement learning improves the overall average to 54.28% while reducing optimal RL steps by 50.7%.
Original abstract
Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process them inefficiently. We present Mage-VL, an efficient codec-native streaming foundation model for real-time multimodal understanding and interaction. At its core, our custom tokenizer, Mage-ViT, replaces uniform frame sampling by selectively encoding dynamic, entropy-rich regions using motion vectors and residual energy across sparse anchor (I) and predicted (P) frames. Operating at a 16 x 16 patch level, this reduces visual token consumption by over 75% while preserving spatiotemporal context. Trained from scratch on approximately 560M unlabeled images and 100M unlabeled video frames, Mage-ViT matches or outperforms flagship encoders trained on billions of image-text pairs. We establish AI4AI data pipelines encompassing prompt-code joint optimization for multimodal captioning and AI-driven performance diagnosis to guide training recipes. Furthermore, through a bio-inspired dual-system architecture - a lightweight System 1 event gate and a causal System 2 decoder - Mage-VL enables proactive streaming perception. Extensive evaluations show that Mage-VL-4B matches Qwen3-VL-4B on static tasks while achieving strong gains in video understanding and 2D/3D spatial reasoning, with up to a 3.5x wall-clock inference speedup, and comprehensively surpasses the 15B Phi-4-reasoning-vision baseline. Beyond model artifacts, we deliver seven key empirical findings covering pre-training data efficiency, variable-resolution scaling, codec system acceleration, VideoQA SFT redundancy, motion-spatial synergy, AI4AI data pipelines, and Zero-Vision SFT for multimodal RL.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.