NTH

VibeVoice-ASR-Streaming Technical Report

AuthorsYujie Tu, Zhiliang Peng, Jianwei Yu, Li Dong, Songchen Xu, Yaoyao Chang, Wenhui Wang, Zilong Wang, Zehua Wang, Yan Xia, Jiajun Zhang, Xie Chen, Furu Wei

September 4, 2026 2 min read
Watch on YouTube
The one-line take

VibeVoice-ASR-Streaming enables an LLM to transcribe and identify speakers in real time, bringing unified speaker-aware speech recognition closer to voice-agent applications.

Key results

24.66
7B five-set recognition mean

Average WER/CER across AliMeeting, AISHELL-4, AMI-SDM, AMI-IHM, and MLC-Challenge.

12 of 13
Speaker-attribution best or tied-best settings

Evaluation settings where the 7B model achieves the best or tied-best cpWER/cpCER.

2.00 s
Expected attribution latency

Steady-state latency for the 7B model with 2.9-second chunks and 0.5-second lookahead.

50,884
Augmented training recordings

Synthetic and real multi-speaker recordings used in training.

4,519.6 hours
Augmented training speech

Total duration of the multi-speaker training mixture.

0.104
Maximum real-time factor

Measured upper-bound serving real-time factor for the 7B 15-frame configuration.

What the paper found

VibeVoice-ASR-Streaming is Microsoft’s end-to-end solution for streaming speaker-attributed automatic speech recognition: instead of running ASR and diarization as separate stages, it interleaves fixed audio chunks, 0.5 seconds of lookahead, and previously generated speaker-labeled text inside a Qwen2.5 language-model context. The released 1.5B and 7B models assign persistent speaker labels while speech arrives, with the 7B configuration using 2.9-second chunks and an expected attribution latency of 2.00 seconds. Across AliMeeting, AISHELL-4, AMI-SDM, AMI-IHM, and the nine-language MLC-Challenge evaluation, it records a five-set recognition mean of 24.66 WER/CER, outperforming deployed systems including OpenAI’s GPT Realtime Whisper, GPT Live Transcribe, ElevenLabs Scribe v2 Realtime, and Gemini 3.5 Transcribe Live, which scores 25.23. For speaker attribution, it achieves the best or tied-best cpWER/cpCER on 12 of 13 settings, while requiring no language prior, speaker count, or external diarization module. Training combines 50,884 augmented multi-speaker recordings totaling 4,519.6 hours, with alignments produced by Qwen3-ForcedAligner-0.6B, followed by offline training, streaming pre-training on roughly 420,000 hours, and curated streaming fine-tuning. Longer chunks and larger models primarily improve speaker attribution, and the 7B system remains real-time, with a measured real-time factor at or below 0.104.

Original abstract

Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-to-end models such as VibeVoice-ASR have unified the two tasks within a single model. However, existing unified models still mainly support offline recognition, making it difficult to meet the low-latency requirements of real-time voice assistants and agents. To tackle this issue, we present VibeVoice-ASR-Streaming, one of the first LLM-based end-to-end approaches to streaming speaker-attributed ASR. It interleaves fixed-size audio chunks, a small amount of lookahead audio and previous text. This allows the model to produce ''who said what'' as speech arrives, without a separate diarization stage. For transcription accuracy, our 7B model achieves the lowest average WER/CER across five evaluation sets. For speaker attribution, it achieves the best or tied-best on 12 of 13 evaluation settings. We release the 1.5B and 7B model weights together with inference code.

Read the original paper

More in Speech AI

Browse all 27 papers →
01Speech

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen

A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.

Read analysis