Qwen3.8-Omni: Towards Native Omni-Modal Agents
AuthorsQwen Team
Resources
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
Key results
Maximum multimodal context length for long-form reasoning and planning
Approximate token volume used in the general multimodal pretraining stage
More than this improvement over Qwen3.5-Omni-Plus across 29 evaluations
Reduction in tokens per query under agentic evidence gathering
Character error rate after autonomous improvement, down from 25.79%
What the paper found
Qwen3.8-Omni-Flash is a native omni-modal agent that jointly reasons over text, images, audio, spatial audio, and video, targeting workflows that systems from OpenAI, Anthropic, DeepSeek, and Google’s Gemini typically address through separate perception and agent layers. Its architecture combines a sparse Mixture-of-Experts backbone inherited from Qwen3.8-Next, Gated DeltaNet, Qwen Sparse Attention, and dedicated AuT and Spatial AuT encoders, with a context window of 1M tokens. Training uses native multimodal co-training, multi-teacher distillation, and reinforcement learning over approximately 2.5 trillion tokens, transferring text-domain planning and tool-use abilities into audiovisual tasks. Against Qwen3.5-Omni-Plus, the model raises average performance by more than 25% across 29 audio, audiovisual, and agent evaluations, while estimated API input costs fall by more than 98% for audio and 93% for audiovisual content. Its agentic evidence-gathering mode improves OmniVideoBench from 63.4 to 67.8 and reduces tokens per query from 145,736 to 79,117, a 45.7% reduction. The accompanying Qwen-MM-Plugins and Qwen-Live-Harness frameworks add selective video retrieval, memory, sub-agent delegation, asynchronous tools, and real-time speech interaction, supporting video editing, translation, meeting execution, research reports, and reusable skills. In an autoresearch demonstration, the model reduced Qwen2.5-Omni-3B’s character error rate on WenetSpeech-Chuan from 25.79% to 15.30%, showing that it can autonomously diagnose multimodal failures and generate targeted training data.
Original abstract
We introduce Qwen3.8-Omni-Flash, a natively multimodal agentic model for real-world multimodal productivity. Compared with previous omni models, which primarily emphasized perception and interaction, Qwen3.8-Omni-Flash substantially improves multimodal understanding and reasoning, as well as performance on long-horizon agentic tasks. These capabilities are supported by a native multimodal co-training strategy that preserves strong text-domain capabilities while facilitating the transfer of agentic capabilities from text to audio and video tasks. The model inherits the sparse mixture-of-experts (MoE) architecture of Qwen3.8-Next and extends the context window to one million tokens, supporting long-context multimodal reasoning and long-horizon planning. These advances enable integration into production workflows as a primary agent or a specialized sub-agent, supporting video editing, long-form audio and video translation, music-conditioned music video or movie generation, and video-based note or omni-skill creation. To address the lack of native audio and video support in existing agent harnesses, we release Qwen-MM-Plugins, a lightweight open-source plugin framework for multimodal productivity. We further frame real-time multimodal interaction as a system-level challenge requiring orchestration of context and memory management, tool use, and sub-agent delegation. Accordingly, we release Qwen-Live-Harness, an open-source framework for building responsive, real-time multimodal agents based on Qwen3.8-Omni-Flash. Extensive evaluations demonstrate that Qwen3.8-Omni-Flash achieves strong performance across multimodal understanding, reasoning, long-horizon agentic execution, and video productivity tasks. These results and the accompanying open-source tools support Qwen3.8-Omni-Flash as a practical foundation for deploying natively multimodal agents in research and production.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.
ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs
Sebastián Andrés Cajas Ordóñez, Maximin Lange, Quang Bui, Anqi Peter Li, Felipe Ocampo Osorio, Rafi Al Attrach, Kushul Reddy Palakala, Sahil Kapadia, Zakaria Laouabdia Sellami, Xinyue Zhang, Ashley Zhang, Leo Anthony Celi
ModaLens tests whether medical VLMs truly look at the X-ray when they already have the report, revealing that reports can substantially suppress measurable image sensitivity.