StepAudio 3 Realtime Technical Report
AuthorsBin Lin, Bo Zhao, Boyang Zhang, Boyong Wu, Chao Yan, Chen Geng, Chen Wu, Cheng Yi, Chengli Feng, Chenglin Zhu, Chengting Feng, Chengyuan Yao, Daijiao Liu, DanNi Wan, Daxin Jiang, Dongjian Li, Dongqing Pang, Fei Tian, Feng Tian, Future Li, Gang Yu, Guanglong Yang, Haoyang Zhang, Hongyuan Wang, Jia Peng, Jiahao Song, Jialong Xue, Jiamin Fan, Jiangjie Zhen, Jianzheng Gao, Jincheng Wen, Jinghua Liang, Jinglan Gong, Jun Chen, Li Xie, Liang Zhao, Lifang Zhang, Lingli Ji, Lun Cai, Min Xu, Peilin Li, Peng Yang, Pengfei Tan, Qingjian Lin, Qinxin Du, Ruijie Xiong, Runze Li, Shenghua Hu, Shengqian Qin, Shi Qiu, Siqi Tu, Siyi Zhou, Tianjiao Deng, Wanying Lu, Weiming Niu, Wen Sun, WenWen Qu, Xiangyu Zhang, Xianwei Zhang, Xiaosu Su, Xing Chen, Xinyu Liu, Xuerui Yang, Yan Wu, Yang Li, Yang Yang, Yechang Huang, Yibo Zhu, Yifan Zhang, Yinuo Yan, Youjun Chen, Yu Fu, Yu Luo, Yu Zhou, Yujie Chen, Yumang Wang, Yunzhou Ju, Yuxiang Yang, Yuxin Li, Yuxin Zhang, Zekai Liu, Zengwei Yao, Zhaoxin Yuan, Zhenwei...
Resources
StepAudio 3 is a real-time voice model that listens, reasons, speaks, and uses tools concurrently for more natural spoken interaction.
Key results
Total parameters in the mixture-of-experts architecture.
Tokens processed across the three pretraining stages.
Audio-understanding score on the MMSU benchmark.
Overall result on the Artificial Analysis Full-Duplex Bench.
Macro task-success rate across airline, retail, and telecom domains.
Speedup from MTP5 with Medusa-style acceptance during private reasoning.
What the paper found
StepAudio 3 Realtime is an audio-language foundation model built around a continuous listen, converse, think, and act loop. Its Deep Perception module combines speech content with paralinguistic and environmental cues, while Seamless Duplex synchronizes user and model audio to handle pauses, backchannels, interruptions, and turn yielding. The model uses a mixture-of-experts design with 196B total parameters and 11B active per token, a Step 3.7 Flash language backbone, and the Audio Transformer encoder from Qwen3-Omni; pretraining processes 1.2T tokens and midtraining extends context to 128K. Its key reasoning innovation, Think-While-Speaking, runs private formulation in parallel with spoken articulation through dual model processes, while Adaptive Thinking selectively invokes deliberation and Medusa-style multi-token prediction accelerates it, reaching a 2.05 wall-clock speedup. In evaluation, reasoning mode scores 73.0 on StepAudioChat, and audio understanding reaches 90.6 on MMSU, exceeding Gemini 3.1 Pro and other baselines on that benchmark. Full-duplex control achieves 98.9 on the Artificial Analysis Full-Duplex Bench, ahead of OpenAI’s GPT-realtime-2 and Qwen Audio 3.0 Realtime Plus. Its asynchronous Voice Agent supports tool execution during ongoing dialogue, achieving 56.0% macro task success on τ-Voice, close to Grok Voice Think Fast 2.0, although retail tasks and multi-turn constraint tracking remain weaknesses.
Original abstract
Realtime spoken interaction demands deep reasoning, prompt responses, and fluid turn-taking. We present StepAudio 3 Realtime, an audio-language foundation model organized around a continuous listen-converse-think-act loop. Deep Perception captures rich acoustic cues to interpret user intent, while Seamless Duplex models synchronized audio streams to handle pauses, backchannels, and interruptions naturally. Crucially, we resolve the tension between deep deliberation and latency via Think-While-Speaking, executing private reasoning in parallel with spoken delivery. In reasoning mode, StepAudio 3 reaches a 73.0 macro average on StepAudioChat. With Think-While-Speaking, it achieves dialogue and reasoning performance comparable to dedicated reasoning models while speaking in real time. Furthermore, an integrated Voice Agent handles asynchronous tool execution without disrupting the dialogue flow. StepAudio 3 Realtime achieves top-tier performance across key dimensions: an exceptional 90.6 on the MMSU benchmark, 98.9 Overall on the Artificial Analysis Full-Duplex Bench, and a 56.0% macro task-success rate on $τ$-Voice.
Read the original paperMore in Speech AI
Browse all 27 papers →Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen
A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.
Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry
Szu-Chi Chen, Jia-Kai Dong, Yi-Cheng Lin, Sung-Feng Huang, Hung-yi Lee
The study argues that voice-similarity systems should be judged by whether their embedding geometry matches human perception, not merely by verification accuracy.
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Jian Chen, You Zhang, Mark Vinton
Tacit-TTS makes zero-shot voice cloning over ten times faster while preserving the ability to clone voices from speech without transcripts, including multilingual and non-lexical references.