NTH

Qwen3.8-Omni: Towards Native Omni-Modal Agents

AuthorsQwen Team

October 2, 2026 2 min read
Watch on YouTube
The one-line take

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Key results

1M
Context window

Maximum multimodal context length for long-form reasoning and planning

2.5T
Pretraining data

Approximate token volume used in the general multimodal pretraining stage

25%
Average evaluation improvement

More than this improvement over Qwen3.5-Omni-Plus across 29 evaluations

45.7%
OmniVideoBench token reduction

Reduction in tokens per query under agentic evidence gathering

15.30%
WenetSpeech-Chuan CER

Character error rate after autonomous improvement, down from 25.79%

What the paper found

Qwen3.8-Omni-Flash is a native omni-modal agent that jointly reasons over text, images, audio, spatial audio, and video, targeting workflows that systems from OpenAI, Anthropic, DeepSeek, and Google’s Gemini typically address through separate perception and agent layers. Its architecture combines a sparse Mixture-of-Experts backbone inherited from Qwen3.8-Next, Gated DeltaNet, Qwen Sparse Attention, and dedicated AuT and Spatial AuT encoders, with a context window of 1M tokens. Training uses native multimodal co-training, multi-teacher distillation, and reinforcement learning over approximately 2.5 trillion tokens, transferring text-domain planning and tool-use abilities into audiovisual tasks. Against Qwen3.5-Omni-Plus, the model raises average performance by more than 25% across 29 audio, audiovisual, and agent evaluations, while estimated API input costs fall by more than 98% for audio and 93% for audiovisual content. Its agentic evidence-gathering mode improves OmniVideoBench from 63.4 to 67.8 and reduces tokens per query from 145,736 to 79,117, a 45.7% reduction. The accompanying Qwen-MM-Plugins and Qwen-Live-Harness frameworks add selective video retrieval, memory, sub-agent delegation, asynchronous tools, and real-time speech interaction, supporting video editing, translation, meeting execution, research reports, and reusable skills. In an autoresearch demonstration, the model reduced Qwen2.5-Omni-3B’s character error rate on WenetSpeech-Chuan from 25.79% to 15.30%, showing that it can autonomously diagnose multimodal failures and generate targeted training data.

Original abstract

We introduce Qwen3.8-Omni-Flash, a natively multimodal agentic model for real-world multimodal productivity. Compared with previous omni models, which primarily emphasized perception and interaction, Qwen3.8-Omni-Flash substantially improves multimodal understanding and reasoning, as well as performance on long-horizon agentic tasks. These capabilities are supported by a native multimodal co-training strategy that preserves strong text-domain capabilities while facilitating the transfer of agentic capabilities from text to audio and video tasks. The model inherits the sparse mixture-of-experts (MoE) architecture of Qwen3.8-Next and extends the context window to one million tokens, supporting long-context multimodal reasoning and long-horizon planning. These advances enable integration into production workflows as a primary agent or a specialized sub-agent, supporting video editing, long-form audio and video translation, music-conditioned music video or movie generation, and video-based note or omni-skill creation. To address the lack of native audio and video support in existing agent harnesses, we release Qwen-MM-Plugins, a lightweight open-source plugin framework for multimodal productivity. We further frame real-time multimodal interaction as a system-level challenge requiring orchestration of context and memory management, tool use, and sub-agent delegation. Accordingly, we release Qwen-Live-Harness, an open-source framework for building responsive, real-time multimodal agents based on Qwen3.8-Omni-Flash. Extensive evaluations demonstrate that Qwen3.8-Omni-Flash achieves strong performance across multimodal understanding, reasoning, long-horizon agentic execution, and video productivity tasks. These results and the accompanying open-source tools support Qwen3.8-Omni-Flash as a practical foundation for deploying natively multimodal agents in research and production.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis
03Multimodal

ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs

Sebastián Andrés Cajas Ordóñez, Maximin Lange, Quang Bui, Anqi Peter Li, Felipe Ocampo Osorio, Rafi Al Attrach, Kushul Reddy Palakala, Sahil Kapadia, Zakaria Laouabdia Sellami, Xinyue Zhang, Ashley Zhang, Leo Anthony Celi

ModaLens tests whether medical VLMs truly look at the X-ray when they already have the report, revealing that reports can substantially suppress measurable image sensitivity.

Read analysis