NTH

Training-Free Speech-Centric Omni Understanding with Frozen VLMs

AuthorsAnkan Deria, Hanoona Rasheed, Xilin He, Fahad Shahbaz Khan, Salman Khan

September 12, 2026 2 min read
Watch on YouTube
The one-line take

This work shows that frozen vision-language models can gain strong multilingual speech and audio-visual abilities simply by connecting them to Whisper, without retraining the models themselves.

Key results

56
Core benchmark datasets

Total distinct datasets used for the main evaluation.

21
Multilingual evaluation languages

CoVoST2 languages used to test multilingual speech understanding.

0.65
Whisper confidence threshold

Minimum segment confidence required before transcript routing.

3.0
Qwen2.5-VL 3B audio-visual gain

Average-point improvement of TFO over native Qwen2.5 Omni at 3B.

18.4
MiniCPM4.5 multilingual gain

Average-point improvement on CoVoST2, from 45.6 to 64.0.

What the paper found

Training-Free Omni, or TFO, turns a frozen vision-language model into a speech-centric omni system without changing its architecture, weights, visual encoder, or requiring audio-video-text training. It uses Whisper-large-v3-turbo to produce confidence-filtered, timestamped transcripts, inserts them into the model’s existing language prompt, and can synthesize replies with CosyVoice3. Across 56 benchmarks and 21 languages, TFO was compared with native omni counterparts built around Qwen2.5-VL, MiniCPM-V, NVILA, and Qwen3-VL, including the Gemini-labeled AVUT-Gemini benchmark and OpenAI’s GPT-5.6 as an evaluator. Speech routing was especially effective when audio evidence was linguistic: Qwen2.5-VL-TFO improved the audio-visual average by 3.0 points at 3B parameters, while audio-only averages improved in all five matched model settings. On multilingual CoVoST2, the MiniCPM4.5 conversion rose from 45.6 to 64.0 across 21 languages, a 18.4-point gain. The method also generally preserved or improved image and video understanding, coding, mathematical reasoning, medical question answering, and visual grounding relative to native omni checkpoints. TFO applies a 0.65 Whisper confidence threshold and omits unreliable audio segments, but its sequential ASR stage increases inference latency and transcripts cannot capture music, environmental sounds, vocal tone, or other non-speech acoustics. The results position modular ASR-to-language routing as a practical alternative to repeatedly retraining native omni models, while showing that dedicated acoustic representations remain necessary for sound-centric reasoning.

Original abstract

Audio-visual understanding remains challenging because models must jointly interpret spoken content, visual events, and their temporal relationships. Existing omni models typically introduce dedicated audio encoders and rely on expensive audio-video-text training, tightly coupling omni capability to specific VLM backbones and potentially weakening their existing visual and reasoning abilities. This raises three questions: whether native omni training is necessary for every new VLM, whether speech-centric omni capability can be added while preserving the original backbone, and where richer acoustic representations remain essential. We introduce Training-Free Omni (TFO), a plug-and-play framework that converts a frozen VLM into a speech-centric omni model without architectural modification, or multimodal re-alignment. TFO uses Whisper to extract confidence-filtered, timestamped transcripts and routes them through the VLM's existing language interface, while leaving its visual pathway unchanged. Across matched comparisons with native omni models on 56 benchmarks and 21 languages, TFO is competitive on audio-visual understanding, improves average audio-only performance across all five model settings, and achieves substantial multilingual speech gains. Freezing the VLM also generally preserves stronger image/video understanding, visual grounding, coding, mathematical reasoning, and medical question answering than the corresponding native omni checkpoints. These results show that strong speech-centric omni understanding can often be obtained through modular audio-to-language routing rather than costly backbone-specific training.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis