NTH

StepAudio 3 Realtime Technical Report

AuthorsBin Lin, Bo Zhao, Boyang Zhang, Boyong Wu, Chao Yan, Chen Geng, Chen Wu, Cheng Yi, Chengli Feng, Chenglin Zhu, Chengting Feng, Chengyuan Yao, Daijiao Liu, DanNi Wan, Daxin Jiang, Dongjian Li, Dongqing Pang, Fei Tian, Feng Tian, Future Li, Gang Yu, Guanglong Yang, Haoyang Zhang, Hongyuan Wang, Jia Peng, Jiahao Song, Jialong Xue, Jiamin Fan, Jiangjie Zhen, Jianzheng Gao, Jincheng Wen, Jinghua Liang, Jinglan Gong, Jun Chen, Li Xie, Liang Zhao, Lifang Zhang, Lingli Ji, Lun Cai, Min Xu, Peilin Li, Peng Yang, Pengfei Tan, Qingjian Lin, Qinxin Du, Ruijie Xiong, Runze Li, Shenghua Hu, Shengqian Qin, Shi Qiu, Siqi Tu, Siyi Zhou, Tianjiao Deng, Wanying Lu, Weiming Niu, Wen Sun, WenWen Qu, Xiangyu Zhang, Xianwei Zhang, Xiaosu Su, Xing Chen, Xinyu Liu, Xuerui Yang, Yan Wu, Yang Li, Yang Yang, Yechang Huang, Yibo Zhu, Yifan Zhang, Yinuo Yan, Youjun Chen, Yu Fu, Yu Luo, Yu Zhou, Yujie Chen, Yumang Wang, Yunzhou Ju, Yuxiang Yang, Yuxin Li, Yuxin Zhang, Zekai Liu, Zengwei Yao, Zhaoxin Yuan, Zhenwei...

September 20, 2026 2 min read
Watch on YouTube
The one-line take

StepAudio 3 is a real-time voice model that listens, reasons, speaks, and uses tools concurrently for more natural spoken interaction.

Key results

196B
Model parameters

Total parameters in the mixture-of-experts architecture.

1.2T
Pretraining tokens

Tokens processed across the three pretraining stages.

90.6
MMSU score

Audio-understanding score on the MMSU benchmark.

98.9
Full-duplex score

Overall result on the Artificial Analysis Full-Duplex Bench.

56.0%
τ-Voice macro task success

Macro task-success rate across airline, retail, and telecom domains.

2.05
MTP wall-clock speedup

Speedup from MTP5 with Medusa-style acceptance during private reasoning.

What the paper found

StepAudio 3 Realtime is an audio-language foundation model built around a continuous listen, converse, think, and act loop. Its Deep Perception module combines speech content with paralinguistic and environmental cues, while Seamless Duplex synchronizes user and model audio to handle pauses, backchannels, interruptions, and turn yielding. The model uses a mixture-of-experts design with 196B total parameters and 11B active per token, a Step 3.7 Flash language backbone, and the Audio Transformer encoder from Qwen3-Omni; pretraining processes 1.2T tokens and midtraining extends context to 128K. Its key reasoning innovation, Think-While-Speaking, runs private formulation in parallel with spoken articulation through dual model processes, while Adaptive Thinking selectively invokes deliberation and Medusa-style multi-token prediction accelerates it, reaching a 2.05 wall-clock speedup. In evaluation, reasoning mode scores 73.0 on StepAudioChat, and audio understanding reaches 90.6 on MMSU, exceeding Gemini 3.1 Pro and other baselines on that benchmark. Full-duplex control achieves 98.9 on the Artificial Analysis Full-Duplex Bench, ahead of OpenAI’s GPT-realtime-2 and Qwen Audio 3.0 Realtime Plus. Its asynchronous Voice Agent supports tool execution during ongoing dialogue, achieving 56.0% macro task success on τ-Voice, close to Grok Voice Think Fast 2.0, although retail tasks and multi-turn constraint tracking remain weaknesses.

Original abstract

Realtime spoken interaction demands deep reasoning, prompt responses, and fluid turn-taking. We present StepAudio 3 Realtime, an audio-language foundation model organized around a continuous listen-converse-think-act loop. Deep Perception captures rich acoustic cues to interpret user intent, while Seamless Duplex models synchronized audio streams to handle pauses, backchannels, and interruptions naturally. Crucially, we resolve the tension between deep deliberation and latency via Think-While-Speaking, executing private reasoning in parallel with spoken delivery. In reasoning mode, StepAudio 3 reaches a 73.0 macro average on StepAudioChat. With Think-While-Speaking, it achieves dialogue and reasoning performance comparable to dedicated reasoning models while speaking in real time. Furthermore, an integrated Voice Agent handles asynchronous tool execution without disrupting the dialogue flow. StepAudio 3 Realtime achieves top-tier performance across key dimensions: an exceptional 90.6 on the MMSU benchmark, 98.9 Overall on the Artificial Analysis Full-Duplex Bench, and a 56.0% macro task-success rate on $τ$-Voice.

Read the original paper

More in Speech AI

Browse all 27 papers →
01Speech

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen

A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.

Read analysis