NTH

VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction

AuthorsZhifei Xie, Jiaqi Lang, Ze An, Yifan Zhao, Dongchao Yang, Kai Li, Ziyang Ma, Mingbao Lin, Chunyan Miao, Shuicheng Yan

September 4, 2026 2 min read
Watch on YouTube
The one-line take

VoiceMem gives real-time voice assistants separate factual and emotional memories so they can respond faster, more personally, and with greater empathy.

Key results

91.2
LoCoMo score

VoiceMem score using top-5 retrieval.

430
Memory budget

Memory tokens used for the 91.2 LoCoMo score.

134
Retrieval latency

Milliseconds required for dense dual-brain retrieval.

1.89
Persona improvement

Aggregate-score improvement over the previous best persona-memory system.

316
ChatMem-Bench questions

Questions testing long-horizon audio memory.

53
ChatMem-Bench audio

Hours of dialogue represented in the benchmark.

What the paper found

VoiceMem proposes a streaming dual-brain memory architecture for real-time voice agents: a left brain organizes factual information through schema–entity graphs and emergent clustering, while a right brain models stable persona traits and context-linked emotions using short- and long-horizon affective attribution. Its four-stage query pipeline processes partial speech, speaker identity, entities, emotion, and embeddings while the user is still talking, then performs compact joint retrieval with a top-5 budget. Using Mem0 as an interchangeable backend, VoiceMem reaches a 91.2 score on LoCoMo with only 430 memory tokens and completes retrieval in 134 milliseconds, keeping memory access inside the latency window of voice activity detection. On persona benchmarks, it improves the aggregate score by 1.89 points over the previous best system. The multimodal ChatMem-Bench contains 316 questions from 53 hours of dialogue and tests information recall, persona reasoning, affective attribution, paralinguistic cues, and environmental sounds. For model adaptation, the system introduces SLM-verified black-box online policy distillation and trains Qwen2.5-Omni, Qwen3-Omni, and Step-Audio2-Mini using the ChatMem-400K corpus; evaluations use GPT-4o-mini and text-embedding-3-small from OpenAI for baseline comparisons. The central result is that dense upper-level routing, rather than returning more memories, enables accurate, emotionally grounded recall at speech-interaction latency.

Original abstract

Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empathetic memory system as their soul. We introduce VoiceMem, a simple memory architecture with a parallel informational left brain, an emotional right brain, and streaming memory I/O mechanisms. We further build a complete pipeline for memory-aware SLM training, long-horizon evaluation, and decoupled deployment with interchangeable memory backends. Experiments and real-world deployment show three advantages: i) Accuracy: under top-5 retrieval, the left brain outperforms classical systems such as Mem0 at top-200 by nearly 30 points; ii) Emotional & Personal: the right brain, with short- and long-horizon affective attribution and dual-node persona modeling, achieves state-of-the-art performance across three persona benchmarks and improves the aggregate score by 4.29 points over the previous best system; and iii) Real-Time & Cheap: VoiceMem completes retrieval in 134 ms, well within standard VAD latency, adding no extra conversational delay while maintaining high accuracy and low cost. These results show that VoiceMem provides a practical memory foundation for real-time, personalized, and emotionally aware speech interaction.

Read the original paper

More in Speech AI

Browse all 27 papers →
01Speech

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen

A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.

Read analysis