Agent Memory: Characterization and System Implications of Stateful Long-Horizon Workloads
cs.AI
Submitted: 2026-06-04
Updated: 2026-09-22
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLM agents are increasingly deployed on long-horizon tasks requiring sustained reasoning over extended interaction histories.
Terminology
Abstract
LLM agents are increasingly deployed on long-horizon tasks requiring sustained reasoning over extended interaction histories. Realizing this at scale requires agents to persistently store, retrieve, and update their own memory across sessions. A rich ecosystem of agent memory systems has emerged spanning flat retrieval, LLM-mediated extraction, consolidating fact stores, and agentic control flows. Yet, their system-level behavior remains uncharacterized. We present the first systems characterization of agent memory. First, we introduce a system-oriented taxonomy classifying agent memory systems along four axes. Second, we build a phase-aware profiling harness attributing cost to construction, retrieval, and generation. Third, we characterize ten representative systems across two benchmark suites, uncovering how design choices shift cost across the write and read paths. Finally, we derive 10 system recommendations covering construction scheduling, capability floors, amortization via query volume, freshness-latency tradeoffs, and fleet-scale management.
Sources
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
- SimpleMem: Efficient Lifelong Memory for LLM Agents
- Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
- RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval
- FP8 Formats for Deep Learning
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- MemGPT: Towards LLMs as Operating Systems
- Zep: A Temporal Knowledge Graph Architecture for Agent Memory
- MIRIX: Multi-Agent Memory System for LLM-Based Agents
- LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
- Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
- A-MEM: Agentic Memory for LLM Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection