VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
eess.AS, cs.AI, cs.IR, cs.MM, cs.SD
Submitted: 2026-08-26
Updated: 2026-08-26
Comments: 18 pages, 9 figures, 6 tables
Project page: https://xzf-thu.github.io/VoiceMem
License: http://creativecommons.org/licenses/by/4.0/
The gist: Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empathetic memory system as their soul.
Terminology
Abstract
Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empathetic memory system as their soul. We introduce VoiceMem, a simple memory architecture with a parallel informational left brain, an emotional right brain, and streaming memory I/O mechanisms. We further build a complete pipeline for memory-aware SLM training, long-horizon evaluation, and decoupled deployment with interchangeable memory backends. Experiments and real-world deployment show three advantages: i) Accuracy: under top-5 retrieval, the left brain outperforms classical systems such as Mem0 at top-200 by nearly 30 points; ii) Emotional & Personal: the right brain, with short- and long-horizon affective attribution and dual-node persona modeling, achieves state-of-the-art performance across three persona benchmarks and improves the aggregate score by 4.29 points over the previous best system; and iii) Real-Time & Cheap: VoiceMem completes retrieval in 134 ms, well within standard VAD latency, adding no extra conversational delay while maintaining high accuracy and low cost. These results show that VoiceMem provides a practical memory foundation for real-time, personalized, and emotionally aware speech interaction.
Sources
- APEX-MEM: Agentic Semi-Structured Memory with Temporal Reasoning for Long-Term Conversational AI
- Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- SHANKS: Simultaneous Hearing and Thinking for Spoken Language Models
- Agents Thinking Fast and Slow: A Talker-Reasoner Architecture
- Moshi: a speech-text foundation model for real-time dialogue
- EverMemOS: A Self-Organizing Memory Operating System for Structured Long-Horizon Reasoning
- Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale
- MAGMA: A Multi-Graph based Agentic Memory Architecture for AI Agents
- MemOS: An Operating System for Memory-Augmented Generation (MAG) in Large Language Models
- Dynamic Affective Memory Management for Personalized LLM Agents
- Qwen3.5-Omni Technical Report
- Zep: A Temporal Knowledge Graph Architecture for Agent Memory
- Qwen3-ASR Technical Report
- Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents
- Step-Audio 2 Technical Report
- Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models
- Chronological Thinking in Full-Duplex Spoken Dialogue Language Models
- GAM: Hierarchical Graph-based Agentic Memory for LLM Agents
- Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions