NTH

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents

AuthorsYinhao Bai, Jinming Chen, Yafeng Chen, Wei Deng, Boya Dong, Nan Duan, Yu Gu, Weisheng Han, Yankun Huang, Ming Ke, Hao Li, Jingdong Li, Xiangyu Liang, Ning Liu, Yuan Liu, Ji Miao, Jiaqi Wang, Qi Wang, Wenchao Wang, Yuxuan Wang, Zhenfang Wang, Zhangyu Xiao, Chao Xue, Hongfei Xue, Fan Yu, Tianyi Zhang, Yuan Zhang, Yuqi Zhang, Lin Zhu

August 9, 2026 2 min read
Watch on YouTube
The one-line take

JoyAI-Talker is a full-duplex empathetic voice agent that combines language reasoning, expressive speech generation, and real-time interruption handling.

Key results

48.9B
Total model parameters

Total parameters in the JoyAI-LLM Flash sparse Mixture-of-Experts Thinker backbone.

3.28B
Active parameters per token

Parameters activated for each input token by the sparse MoE backbone.

94.62%
MATH score

Text-to-text mathematical reasoning score for JoyAI-Talker-DPO.

0.88
Interruption response rate

CRESPOND rate on genuine user interruptions in Full-Duplex-Bench v1.5.

0.10
Background false-trigger rate

CRESPOND rate under background speech in Full-Duplex-Bench v1.5.

What the paper found

JoyAI-Talker is a full-duplex speech dialogue system designed to combine foundation-model reasoning with empathetic, expressive voice interaction. Its decoupled Duplex-Thinker-Talker architecture separates semantic planning, conversational state control, and speech synthesis: the Thinker uses the 48.9B-parameter JoyAI-LLM Flash sparse Mixture-of-Experts backbone, with 3.28B parameters activated per token, while the Talker generates streaming speech through a Transformer-DiT pipeline. Unlike tightly coupled systems such as Moshi, the model applies unified speech-text joint training from mid-training through supervised fine-tuning and DPO to reduce cognitive degradation, preserving strong textual reasoning; it scores 94.62% on MATH. The Persona-Adaptive Empathetic Response, or PAER, framework extracts speaker attributes and emotional cues from audio, incorporates them into Chain-of-Thought reasoning, and controls vocal delivery through natural-language instructions plus localized tokens such as [Laughter] and [Sigh]. Joy-Duplex adds a lightweight, state-driven semantic gate using interleaved partial, complete, accept, reject, and backchannel tokens, avoiding many false activations associated with energy-based voice activity detection. On Full-Duplex-Bench v1.5, it achieves a 0.88 response rate to genuine user interruptions and a 0.10 false-trigger rate under background speech. The system compares favorably with Qwen3-Omni, Moshi, GPT-4o, and Gemini 3.1 Live, while its modular design trades a small amount of latency and acoustic detail for independent optimization and easier deployment.

Original abstract

We present JoyAI-Talker, a full-duplex speech dialogue system that delivers robust foundation model capabilities while empowering empathetic interaction and voice agent intelligence. JoyAI-Talker adopts a modular Thinker-Talker architecture and further implements a unified speech-text joint training pipeline to mitigate the common "cognitive degradation" bottleneck, thereby largely preserving the model's core textual reasoning, STEM, and logical capabilities while extending them to speech-based interaction. For expressive speech synthesis, the Talker module employs a text-controllable generation paradigm that enables natural-language instructions to flexibly control vocal attributes and localized paralinguistic events, such as laughter and sighs, supporting more expressive and fine-grained speech responses. To enhance conversational empathy, we introduce the Persona-Adaptive Empathetic Response (PAER) framework. PAER employs a hierarchical cognitive pipeline to extract non-verbal speaker cues, such as gender, age, and emotional state, from raw input audio, incorporate them into the Thinker's CoT reasoning, and generate context-adaptive responses that align semantically appropriate text with fine-grained control over utterance-level expressiveness and localized paralinguistic events, including sighs, speaking rate, and volume. We further integrate Joy-Duplex, a state-driven, plug-and-play full-duplex framework that functions as an efficient gating engine for real-time turn control. Extensive evaluations show that JoyAI-Talker achieves highly competitive performance on foundational T2T and S2T benchmarks. In full-duplex evaluation, the system reaches a high response rate of 0.88 under user interruptions while maintaining an extremely low false-trigger rate under background speech, demonstrating its readiness for fluid and natural speech dialogue.

Read the original paper

More in Speech AI

Browse all 27 papers →
01Speech

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen

A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.

Read analysis