StreamMemBench: Streaming Evaluation of Agent Memory for Future-Oriented Assistance
cs.AI
Submitted: 2026-06-12
Updated: 2026-08-26
Comments: Accepted to Findings of EMNLP 2026
Code: https://github.com/landian60/StreamMemBench
License: http://creativecommons.org/licenses/by/4.0/
The gist: A central role of personal-agent memory is to turn stored information and prior interactions into future-oriented assistance.
Terminology
Abstract
A central role of personal-agent memory is to turn stored information and prior interactions into future-oriented assistance. In daily use, useful cues come from what the agent observes and how the user interacts with the agent, and the agent must carry them forward from the current request to similar future tasks. Existing memory benchmarks usually test dialogue recall or task improvement in isolation, leaving the trajectory from streaming observations to later assistance largely untested. We introduce StreamMemBench, a streaming benchmark that constructs a two-step task sequence around each evidence anchor from EgoLife egocentric streams. The initial task tests evidence use, while the follow-up task tests whether feedback and interaction experience are reused. Four metrics diagnose evidence recall, initial evidence use, feedback incorporation, and follow-up reuse. Experiments with eight memory systems across two backbones show that current systems often fail to use observed evidence or turn feedback into reliable follow-up behavior, even when evidence is stored or feedback is incorporated locally. StreamMemBench is publicly available at https://github.com/landian60/StreamMemBench.
Sources
- MemOS: A Memory OS for AI System
- MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems
- IntPro: A Proxy Agent for Context-Aware Intent Understanding via Retrieval-conditioned Inference
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- Memp: Exploring Agent Procedural Memory
- MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks
- Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
- EverMemOS: A Self-Organizing Memory Operating System for Structured Long-Horizon Reasoning
- Evaluating Long-Horizon Memory for Multi-Party Collaborative Dialogues
- MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents
- Evaluating Memory Capability in Continuous Lifelog Scenario
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection