NTH

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

AuthorsSenqiao Yang, Kaichen Zhang, Zhaoyang Jia, Jinghao Guo, Yifei Shen, Xinjie Zhang, Xiaoyi Zhang, Haoqing Wang, Xiao Li, Peng Zhang, Xiang An, Yin Xie, Zhening Liu, Xun Guo, Jiahao Li, Shicheng Zheng, Jinglu Wang, Zongyu Guo, Wenxuan Xie, Zihan Zheng, Yuxuan Luo, Bin Li, Yan Lu

August 3, 2026 3 min read
Watch on YouTube
The one-line take

Mage-VL turns video codecs and motion-aware tokenization into a faster multimodal model designed for real-time visual understanding.

Key results

75%
Visual token reduction

Codec-native patch selection removes redundant visual tokens from streaming video.

560M
Unlabeled image pretraining

Mage-ViT is trained from scratch on unlabeled images.

100M
Unlabeled video-frame pretraining

Mage-ViT is jointly trained on unlabeled video frames.

4B
Mage-VL model size

Mage-VL-4B uses the Qwen3-4B-Instruct language backbone.

3.5
Inference speedup

Codec-native processing achieves up to 3.5× wall-clock speedup over uniform frame sampling.

64.00
OVO-Bench score

Mage-VL-4B achieves the reported overall streaming score on OVO-Bench.

What the paper found

Microsoft’s Mage-VL addresses the streaming weakness of conventional vision-language models by making video compression part of the visual tokenizer. Its Mage-ViT encoder divides frames into 16×16 patches, densely processes anchor I-frames, and selectively retains motion- and entropy-rich patches from predicted P-frames using HEVC/H.265 motion vectors and residual energy; the same strategy also works with the neural codec DCVC-RT. Trained from scratch on 560M unlabeled images and 100M unlabeled video frames, Mage-ViT rivals encoders such as SigLIP2 that rely on far larger image-text corpora. The resulting Mage-VL-4B combines this encoder with the Qwen3-4B-Instruct language backbone and a five-stage curriculum covering captions, instruction tuning, long-video adaptation, codec-native training, and streaming alignment. A lightweight System 1 cognition gate decides when an event merits a response, while a causal System 2 decoder generates commentary only after activation. Mage-VL reduces visual token consumption by 75% and delivers up to 3.5× wall-clock inference speedup over uniform frame sampling, while matching Qwen3-VL-4B on static tasks and improving temporal grounding, video understanding, and spatial reasoning. On OVO-Bench, it reaches 64.00%, and its AI4AI pipeline—using GPT-5 rubric scoring and GitHub Copilot prompt-code refinement—improves caption quality and downstream benchmarks. The paper also reports that Zero-Vision SFT followed by multimodal reinforcement learning improves the overall average to 54.28% while reducing optimal RL steps by 50.7%.

Original abstract

Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process them inefficiently. We present Mage-VL, an efficient codec-native streaming foundation model for real-time multimodal understanding and interaction. At its core, our custom tokenizer, Mage-ViT, replaces uniform frame sampling by selectively encoding dynamic, entropy-rich regions using motion vectors and residual energy across sparse anchor (I) and predicted (P) frames. Operating at a 16 x 16 patch level, this reduces visual token consumption by over 75% while preserving spatiotemporal context. Trained from scratch on approximately 560M unlabeled images and 100M unlabeled video frames, Mage-ViT matches or outperforms flagship encoders trained on billions of image-text pairs. We establish AI4AI data pipelines encompassing prompt-code joint optimization for multimodal captioning and AI-driven performance diagnosis to guide training recipes. Furthermore, through a bio-inspired dual-system architecture - a lightweight System 1 event gate and a causal System 2 decoder - Mage-VL enables proactive streaming perception. Extensive evaluations show that Mage-VL-4B matches Qwen3-VL-4B on static tasks while achieving strong gains in video understanding and 2D/3D spatial reasoning, with up to a 3.5x wall-clock inference speedup, and comprehensively surpasses the 15B Phi-4-reasoning-vision baseline. Beyond model artifacts, we deliver seven key empirical findings covering pre-training data efficiency, variable-resolution scaling, codec system acceleration, VideoQA SFT redundancy, motion-spatial synergy, AI4AI data pipelines, and Zero-Vision SFT for multimodal RL.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis