CMI-Mem: Toward Generalizable Long-Term Memory Management via CMI-Augmented Reinforcement Learning
Yubo Wang, Qiuyu Zhao, Zenghui Sun, Shichao Dong, Jinsong Lan, Xiaoyong Zhu, Haoyang Li, Bo Zheng, Lei Chen
cs.AI, cs.CL
Submitted: 2026-08-22
Updated: 2026-08-25
Code: https://github.com/Wyb0627/CMIMem
License: http://creativecommons.org/licenses/by/4.0/
The gist: Memory Manager models are pivotal in agent systems.
Terminology
Abstract
Memory Manager models are pivotal in agent systems. Existing reinforcement-learning methods commonly use LLM-judged synthetic question-answer (QA) pairs: this provides useful downstream task grounding, but values memory through a sampled query distribution and a fixed reader. We propose CMI-Mem, a lightweight RL memory manager with a hybrid reward. Its extrinsic QA term measures end-task correctness, while its intrinsic Conditional Mutual Information (CMI) term evaluates the information contributed by new conversational inputs relative to the current memory state without conditioning on a sampled QA query. The two signals are complementary: QA anchors task utility, whereas CMI provides per-operation supervision for relevant, non-redundant memory construction. Experiments demonstrate improved transfer across memory-use scenarios, together with more efficient training and inference from the per-operation CMI signal. Our codes are available at: https://github.com/Wyb0627/CMIMem, and the CMI-Mem-4B model checkpoint is available at: https://www.modelscope.cn/models/wyb0627/CMIMem-4B
Sources
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- DeepSeek-V3 Technical Report
- AriGraph: Learning Knowledge Graph World Models with Episodic Memory for LLM Agents
- MemOS: A Memory OS for AI System
- MemBuilder: Reinforcing LLMs for Long-Term Memory Construction via Attributed Dense Rewards
- Regularization and Reparameterization Avoid Vanishing Gradients in Sigmoid-Type Networks
- Fine-Mem: Fine-Grained Feedback Alignment for Long-Horizon Memory Management
- MIRIX: Multi-Agent Memory System for LLM-Based Agents
- Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
- Zep: A Temporal Knowledge Graph Architecture for Agent Memory
- Qwen3 Technical Report
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents
- DeltaMem: Towards Agentic Memory Management via Reinforcement Learning
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
- PersonalLLM: Tailoring LLMs to Individual Preferences
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection