MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks
cs.CL
Submitted: 2026-02-18
Updated: 2026-09-17
Comments: ICML 2026
Project page: https://memoryarena.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems
- Beyond Task Completion: An Assessment Framework for Evaluating Agentic AI Systems
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent
- AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents
- VeriWeb: Verifiable Long-Chain Web Benchmark for Agentic Information-Seeking
- MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory
- AgentLongBench: A Controllable Long Benchmark For Long-Contexts Agents via Environment Rollouts
- Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- Memory in the Age of AI Agents
- Mem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous Agents
- Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
- Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering