NTH

Xiaomi-CocktailASR-1 Technical Report

AuthorsYiru Zhang, Hang Su, Lichun Fan, Ying Zeng, Chang Liu, Yifeng Wang, Yuquan Liang, Tao Li, Lian Li, Wenhao Yang, Jian Luan, Cong Zou, Heng Qu

September 12, 2026 2 min read
Watch on YouTube
The one-line take

Xiaomi-CocktailASR-1 uses a voice sample as a prompt to transcribe only the desired speaker in noisy conversations, while rejecting cases where that speaker is absent.

Key results

0.6B
Audio encoder size

Parameter count of the Data2Vec2-derived audio encoder

400,000
Multi-speaker training data

Approximate hours of real and synthetic multi-speaker mixtures

2.90%
LibriSpeechMix 2mix TS-WER

Target-speaker word error rate on a synthetic overlapping-speech benchmark

20.63%
AliMeeting-Far TS-WER

Target-speaker word error rate on real-world Chinese meeting audio

79.59%
LibriSpeech Neg rejection rate

Correct rejection rate when the target speaker is absent

3.87%
CoT LibriMix 2mix WER

WER with Chain-of-Thought reasoning, versus 4.11% in standard mode

What the paper found

Xiaomi-CocktailASR-1 is an end-to-end target-speaker ASR system designed for the cocktail-party problem: given a one-to-four-second reference recording, it transcribes that speaker directly from mixed audio without explicit speech separation. Its 0.6B-parameter Data2Vec2-derived audio encoder uses FBank features and a reference voiceprint prompt, while an adapter connects speech representations to the Qwen3-8B language model. Training combines approximately 400,000 hours of multi-speaker mixtures, 600,000 hours of single-speaker data, and 10,000 hours of negative samples, using staged supervised fine-tuning, noise mixing, and reinforcement learning. Unlike general ASR models such as Qwen3-ASR-1.7b and StepAudio2, which perform poorly on overlapping speech, the system also rejects inputs where the target is absent and supports Chain-of-Thought reasoning. It reaches a 2.90% TS-WER on LibriSpeechMix 2mix and 20.63% on the real-world AliMeeting-Far benchmark, while maintaining competitive single-speaker accuracy. On LibriSpeech Neg, it achieves a 79.59% rejection rate, and CoT reduces LibriMix 2mix WER from 4.11% to 3.87%. Compared with Gemini-2.5-pro, Xiaomi-CocktailASR-1 is specialized for target-speaker recognition and balances interference suppression, target-present accuracy, and rejection reliability in one architecture.

Original abstract

Recently, large language model (LLM) based ASR models have achieved significant progress, yet they generally lack support for multi-speaker scenarios, where the cocktail party problem remains a critical bottleneck for further advancing ASR. Existing TS-ASR methods, including end-to-end architectures with speaker embeddings and latest LLM-based explorations suffer from degraded single-speaker performance and the inability to reject when the target speaker is absent. In this paper, we propose Xiaomi-CocktailASR-1, an LLM-based end-to-end TS-ASR architecture. By utilizing reference speech as voiceprint prompts, it directly transcribes the target speaker's speech without requiring speech separation. Xiaomi-CocktailASR-1 maintains competitive performance in single-speaker scenarios, comparable to mainstream ASR models. It also features a negative sample rejection capability, outputting empty text when the target speaker is absent from the mixed speech. Additionally, Xiaomi-CocktailASR-1 supports a Chain-of-Thought (CoT) reasoning mode to provide explicit reasoning steps. Extensive experiments on various synthetic and real-world multispeaker benchmarks demonstrate that Xiaomi-CocktailASR-1 achieves state-of-the-art performance, effectively addressing the cocktail party problem through a unified architecture that balances multispeaker and single-speaker recognition accuracy, along with rejection capability.

Read the original paper

More in Speech AI

Browse all 27 papers →
01Speech

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen

A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.

Read analysis