Xiaomi-CocktailASR-1 Technical Report
AuthorsYiru Zhang, Hang Su, Lichun Fan, Ying Zeng, Chang Liu, Yifeng Wang, Yuquan Liang, Tao Li, Lian Li, Wenhao Yang, Jian Luan, Cong Zou, Heng Qu
Resources
Xiaomi-CocktailASR-1 uses a voice sample as a prompt to transcribe only the desired speaker in noisy conversations, while rejecting cases where that speaker is absent.
Key results
Parameter count of the Data2Vec2-derived audio encoder
Approximate hours of real and synthetic multi-speaker mixtures
Target-speaker word error rate on a synthetic overlapping-speech benchmark
Target-speaker word error rate on real-world Chinese meeting audio
Correct rejection rate when the target speaker is absent
WER with Chain-of-Thought reasoning, versus 4.11% in standard mode
What the paper found
Xiaomi-CocktailASR-1 is an end-to-end target-speaker ASR system designed for the cocktail-party problem: given a one-to-four-second reference recording, it transcribes that speaker directly from mixed audio without explicit speech separation. Its 0.6B-parameter Data2Vec2-derived audio encoder uses FBank features and a reference voiceprint prompt, while an adapter connects speech representations to the Qwen3-8B language model. Training combines approximately 400,000 hours of multi-speaker mixtures, 600,000 hours of single-speaker data, and 10,000 hours of negative samples, using staged supervised fine-tuning, noise mixing, and reinforcement learning. Unlike general ASR models such as Qwen3-ASR-1.7b and StepAudio2, which perform poorly on overlapping speech, the system also rejects inputs where the target is absent and supports Chain-of-Thought reasoning. It reaches a 2.90% TS-WER on LibriSpeechMix 2mix and 20.63% on the real-world AliMeeting-Far benchmark, while maintaining competitive single-speaker accuracy. On LibriSpeech Neg, it achieves a 79.59% rejection rate, and CoT reduces LibriMix 2mix WER from 4.11% to 3.87%. Compared with Gemini-2.5-pro, Xiaomi-CocktailASR-1 is specialized for target-speaker recognition and balances interference suppression, target-present accuracy, and rejection reliability in one architecture.
Original abstract
Recently, large language model (LLM) based ASR models have achieved significant progress, yet they generally lack support for multi-speaker scenarios, where the cocktail party problem remains a critical bottleneck for further advancing ASR. Existing TS-ASR methods, including end-to-end architectures with speaker embeddings and latest LLM-based explorations suffer from degraded single-speaker performance and the inability to reject when the target speaker is absent. In this paper, we propose Xiaomi-CocktailASR-1, an LLM-based end-to-end TS-ASR architecture. By utilizing reference speech as voiceprint prompts, it directly transcribes the target speaker's speech without requiring speech separation. Xiaomi-CocktailASR-1 maintains competitive performance in single-speaker scenarios, comparable to mainstream ASR models. It also features a negative sample rejection capability, outputting empty text when the target speaker is absent from the mixed speech. Additionally, Xiaomi-CocktailASR-1 supports a Chain-of-Thought (CoT) reasoning mode to provide explicit reasoning steps. Extensive experiments on various synthetic and real-world multispeaker benchmarks demonstrate that Xiaomi-CocktailASR-1 achieves state-of-the-art performance, effectively addressing the cocktail party problem through a unified architecture that balances multispeaker and single-speaker recognition accuracy, along with rejection capability.
Read the original paperMore in Speech AI
Browse all 27 papers →Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen
A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.
Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry
Szu-Chi Chen, Jia-Kai Dong, Yi-Cheng Lin, Sung-Feng Huang, Hung-yi Lee
The study argues that voice-similarity systems should be judged by whether their embedding geometry matches human perception, not merely by verification accuracy.
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Jian Chen, You Zhang, Mark Vinton
Tacit-TTS makes zero-shot voice cloning over ten times faster while preserving the ability to clone voices from speech without transcripts, including multilingual and non-lexical references.