VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
AuthorsZhifei Xie, Jiaqi Lang, Ze An, Yifan Zhao, Dongchao Yang, Kai Li, Ziyang Ma, Mingbao Lin, Chunyan Miao, Shuicheng Yan
Resources
VoiceMem gives real-time voice assistants separate factual and emotional memories so they can respond faster, more personally, and with greater empathy.
Key results
VoiceMem score using top-5 retrieval.
Memory tokens used for the 91.2 LoCoMo score.
Milliseconds required for dense dual-brain retrieval.
Aggregate-score improvement over the previous best persona-memory system.
Questions testing long-horizon audio memory.
Hours of dialogue represented in the benchmark.
What the paper found
VoiceMem proposes a streaming dual-brain memory architecture for real-time voice agents: a left brain organizes factual information through schema–entity graphs and emergent clustering, while a right brain models stable persona traits and context-linked emotions using short- and long-horizon affective attribution. Its four-stage query pipeline processes partial speech, speaker identity, entities, emotion, and embeddings while the user is still talking, then performs compact joint retrieval with a top-5 budget. Using Mem0 as an interchangeable backend, VoiceMem reaches a 91.2 score on LoCoMo with only 430 memory tokens and completes retrieval in 134 milliseconds, keeping memory access inside the latency window of voice activity detection. On persona benchmarks, it improves the aggregate score by 1.89 points over the previous best system. The multimodal ChatMem-Bench contains 316 questions from 53 hours of dialogue and tests information recall, persona reasoning, affective attribution, paralinguistic cues, and environmental sounds. For model adaptation, the system introduces SLM-verified black-box online policy distillation and trains Qwen2.5-Omni, Qwen3-Omni, and Step-Audio2-Mini using the ChatMem-400K corpus; evaluations use GPT-4o-mini and text-embedding-3-small from OpenAI for baseline comparisons. The central result is that dense upper-level routing, rather than returning more memories, enables accurate, emotionally grounded recall at speech-interaction latency.
Original abstract
Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empathetic memory system as their soul. We introduce VoiceMem, a simple memory architecture with a parallel informational left brain, an emotional right brain, and streaming memory I/O mechanisms. We further build a complete pipeline for memory-aware SLM training, long-horizon evaluation, and decoupled deployment with interchangeable memory backends. Experiments and real-world deployment show three advantages: i) Accuracy: under top-5 retrieval, the left brain outperforms classical systems such as Mem0 at top-200 by nearly 30 points; ii) Emotional & Personal: the right brain, with short- and long-horizon affective attribution and dual-node persona modeling, achieves state-of-the-art performance across three persona benchmarks and improves the aggregate score by 4.29 points over the previous best system; and iii) Real-Time & Cheap: VoiceMem completes retrieval in 134 ms, well within standard VAD latency, adding no extra conversational delay while maintaining high accuracy and low cost. These results show that VoiceMem provides a practical memory foundation for real-time, personalized, and emotionally aware speech interaction.
Read the original paperMore in Speech AI
Browse all 27 papers →Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen
A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.
Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry
Szu-Chi Chen, Jia-Kai Dong, Yi-Cheng Lin, Sung-Feng Huang, Hung-yi Lee
The study argues that voice-similarity systems should be judged by whether their embedding geometry matches human perception, not merely by verification accuracy.
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Jian Chen, You Zhang, Mark Vinton
Tacit-TTS makes zero-shot voice cloning over ten times faster while preserving the ability to clone voices from speech without transcripts, including multilingual and non-lexical references.