Evaluating Memory Structure in LLM Agents
cs.LG, cs.CL
Submitted: 2026-02-11
Updated: 2026-10-01
Comments: Preprint, work in progress
Code: https://github.com/yandex-research/StructMemEval
License: http://creativecommons.org/licenses/by/4.0/
The gist: Modern LLM-based agents and chat assistants rely on long-term memory frameworks to store reusable knowledge, recall user preferences, and augment reasoning.
Terminology
Abstract
Modern LLM-based agents and chat assistants rely on long-term memory frameworks to store reusable knowledge, recall user preferences, and augment reasoning. As researchers create more complex memory architectures, it becomes increasingly difficult to analyze their capabilities and guide future memory designs. Most long-term memory benchmarks focus on simple fact retention, multi-hop recall, and time-based changes. While undoubtedly important, these capabilities can often be achieved with simple retrieval-augmented LLMs and do not test complex memory hierarchies. To bridge this gap, we propose StructMemEval - a benchmark that tests the agent's ability to organize its long-term memory, not just factual recall. We gather a suite of tasks that humans solve by organizing their knowledge in a specific structure: transaction ledgers, to-do lists, trees and others. Our initial experiments show that simple retrieval-augmented LLMs struggle with these tasks, whereas memory agents can reliably solve them if prompted how to organize their memory. However, we also find that modern LLMs do not always recognize the memory structure when not prompted to do so. This highlights an important direction for future improvements in both LLM training and memory frameworks.
Sources
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- SuperLocalMemory: Privacy-Preserving Multi-Agent Memory with Bayesian Trust Defense Against Memory Poisoning
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- Forget What You Know about LLMs Evaluations -- LLMs are Like a Chameleon
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Is Evaluation Awareness Just Format Sensitivity? Limitations of Probe-Based Evidence under Controlled Prompt Structure
- Context Length Alone Hurts LLM Performance Despite Perfect Retrieval
- LightMem: Lightweight and Efficient Memory-Augmented Generation
- ResearchCodeAgent: An LLM Multi-Agent System for Automated Codification of Research Methodologies
- ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
- LLM-Empowered Embodied Agent for Memory-Augmented Task Planning in Household Robotics
- Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
- Memory in the Age of AI Agents
- A Survey of Agent Memory in the Second Half: Towards Self-Evolving and Long-Horizon Agents
- PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
- Memory OS of AI Agent
- DialSim: A Dialogue Simulator for Evaluating Long-Term Multi-Party Dialogue Understanding of Conversational Agents
- CarMem: Enhancing Long-Term Memory in LLM Voice Assistants through Category-Bounding
- Measuring AI Ability to Complete Long Software Tasks
- LLMs Get Lost In Multi-Turn Conversation
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks