DolphinBench: Mapping the Pareto Frontier of Agent Memory
cs.CL, cs.AI
Submitted: 2026-09-21
Updated: 2026-09-22
Comments: 6 pages, 2 figures
Code: https://github.com/anthropics/claude-code
License: http://creativecommons.org/licenses/by/4.0/
The gist: Agents today often take real-world actions that depend on long-term memory and context recall over time.
Terminology
Abstract
Agents today often take real-world actions that depend on long-term memory and context recall over time. However, most current memory benchmarks are built for a conversational question-answer format, where the question itself signals that some fact must be retrieved, and often which one. Moreover, benchmarks rarely require anything beyond accuracy from submissions, allowing memory systems to make unreasonable cost/time tradeoffs to achieve higher scores. We present DolphinBench, a benchmark that evaluates memory directly through an agent's task completion. DolphinBench includes three knowledge-work personas with roughly 500k tokens of user messages per persona and evaluates agents on tasks that depend on information from that history. We verify all 200 tasks per persona by running an agent with and without the relevant history, requiring success with it and failure without it. Finally, we require all evaluations to report total cost and latency alongside accuracy, which enables us to evaluate agent memory systems holistically. No existing memory benchmark combines all three. The dataset and evaluation code are available at https://dolphinbench.ai.
Sources
- MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems
- RealMem: Benchmarking LLMs in Real-World Memory-Driven Interaction
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
- MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks
- Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Hindsight is 20/20: Building Agent Memory that Retains, Recalls, and Reflects
- From Tools to Teammates: Evaluating LLMs in Multi-Session Coding Interactions
- REALTALK: A 21-Day Real-World Dataset for Long-Term Conversation
- Zep: A Temporal Knowledge Graph Architecture for Agent Memory
- Mem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous Agents
- MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents
- MemAudit: Auditing Long-Term Agent Memory via Hidden User-State Recovery
- Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs
- Evaluating Very Long-Term Conversational Memory of LLM Agents
- StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns
- MemGPT: Towards LLMs as Operating Systems
- LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering