MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale

arXiv:2608.02613 · cs.CL, cs.AI, cs.LG, cs.MA · Submitted 2026-05-20 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale".

Tom: As a diligent researcher, I have meticulously analyzed both provided texts regarding "MEMARENA." The goal is to synthesize these two descriptions into a single, comprehensive,

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So, looking at what the authors claim in "MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale," they are setting up a single-world conversational benchmark using the MASIM agent simulator. Their thesis is that existing benchmarks don't adequately test how these agents handle long-horizon, ego-centric context extraction and complex agentic interactions.

Jane: They claim this benchmark fills those gaps by simulating fifty agents interacting over fifteen days, resulting in a massive corpus of ten point three million dialog-text tokens and around 24 point 1K text-only egoobserved tokens per agent per day, which they co-generate ground truth across recall, reasoning, and trustworthiness dimensions based on the interaction history.

Lu: The paper emphasizes that they define six task types spanning those three dimensions—recall, reasoning, and trustworthiness—using user-specific evidence projections to eliminate the need for manual annotation. That approach to creating evaluation instances is quite clever because it scales without needing huge amounts of human labeling upfront.

Meng: I'm interested in how they structured the evaluation around those dimensions; specifically, how they tied those six task types directly back to the core memory functions we care about on a practical level <ref:2608.02613#pg2>.

Lalam: From my perspective, this structure is vital because it forces us to look beyond just whether an answer is correct and instead assess the agent's ability to actually use that information reliably in context.

Tom: And they test five different memory backends—Vanilla context, BM25-RAG, Oracle retrieval, Memobase, and MemSearch—against five open-weight reader models to see which backend truly matters for content accuracy.

Conclusion: Tom: So, wrapping up our look at "MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale," the authors have presented a comprehensive way to stress-test personal memory agents operating on devices. They've shown that the choice of backend matters significantly for content accuracy, with MemSearch showing gains over Memobase, and they've highlighted a universal failure point regarding permission-aware access enforcement where Oracle retrieval leaks information heavily.

Jane: The authors are really pointing toward the critical importance of evidence selection and access control rather than just relying on simple retrieval rankings or increasing context volume to improve performance. This suggests that designing robust systems around how they manage and disclose private data is a bigger hurdle than just making sure they find the right text in the first place.

Lu: The implication for future research, I think, is that we need to focus less on tweaking the retrieval mechanism itself and more on establishing solid protocols for provenance tokens and access metadata to ensure privacy and accuracy simultaneously across different agents.

Meng: If they are correct about the failure of permission-aware access being universal, it means any system we build needs a very strict, almost hardware-level enforcement layer when dealing with sensitive personal memories. I'm thinking about how that translates into actual deployment constraints.

Lalam: For the cultural impact, I see this as moving us toward an era where personal AI assistants can be trusted not just for answering questions, but for managing complex relationships and respecting privacy boundaries in a way we haven't seen before.

Tom: It's definitely a lot to digest, showing us exactly what we need to evaluate when developing these next generation on-device memory systems.

MBZUAI

cs.CL, cs.AI, cs.LG, cs.MA

Submitted: 2026-05-20

Updated: 2026-10-02

Comments: Accepted at NeurIPS 2026 (Evaluations & Datasets Track). 53 pages. Code: https://github.com/dereksodo/MemArena-Bench; data: https://huggingface.co/datasets/dereksodo/memarena-l

Code: https://github.com/memodb-io/memobase

Project page: https://langchain-ai.github.io/langmem

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: As a diligent researcher, I have meticulously analyzed both provided texts regarding "MEMARENA." The goal is to synthesize these two descriptions into a single, comprehensive, and highly detailed

Key concepts

Ego-centric Benchmark
A testing framework specifically designed to measure how well personal memory assistants handle information centered around a single user's history and context. It focuses on the agent's ability to remember, reason about, and safely use that specific user data during interactions.
Memory Backends
Different software systems used to store and retrieve stored information for an agent. The study tested five distinct methods: Vanilla context, BM25-RAG, Oracle retrieval, Memobase, and MemSearch. The goal was to see which storage method provides the most accurate information when the agent needs to recall something.
Trustworthiness
A metric used to judge how reliable and safe an agent's response is. This involves checking if the agent correctly uses its memory according to rules, such as access control policies, ensuring it doesn't leak sensitive data or provide incorrect answers based on its stored facts.

Terminology

Summary

As a diligent researcher, I have meticulously analyzed both provided texts regarding MEMARENA. The goal is to synthesize these two descriptions into a single, comprehensive, and highly detailed summary that captures all critical aspects of the research for maximum clarity and accuracy.

Here is the combined, in-depth summary:


The paper introduces MEMARENA, a novel, ego-centric benchmark meticulously designed to evaluate the performance of personal, on-device memory agents. This benchmark addresses critical gaps in existing memory evaluation methods by focusing specifically on how these agents handle long-horizon, ego-centric context extraction and complex agentic interactions.

MEMARENA is structured as a single-world conversational benchmark built using the MASIM agent simulator. The scale of the evaluation is substantial: it simulates interactions for 50 agents over 15 days, involving a massive corpus of 10.3 million dialog-text tokens and approximately 24.1K text-only egoobserved tokens per agent per day.

The core focus of the evaluation is assessing three fundamental dimensions of memory agent performance:

  1. Memory Recall: The ability to accurately retrieve stored information.

  2. Memory Reasoning: The capacity to use retrieved context to derive logical conclusions or perform complex tasks.

  3. Trustworthiness: Evaluating the reliability and safety of the agent's responses based on its memory utilization and access control enforcement.

The evaluation protocol is rigorous, utilizing a fixed set of 90 queries selected by sub-dimension, with 10 latest queries per sub-dimension timestamped by metadata to ensure temporal accuracy. The study employs a single-user, single-stream latency and energy methodology executed on an NVIDIA GB10 Spark edge node to provide faithful estimates of the end-user experience.

The benchmark systematically evaluates the interplay between different components:

  • Memory Backends: Five comparable memory backends are tested: Vanilla context, BM25-RAG, Oracle retrieval, Memobase, and MemSearch.

  • Answer Models (Readers): The system is paired with various answer models across different reader tiers. Evaluation involves five open-weight readers.

  • Sub-Dimensions: Performance is measured across nine detailed sub-dimensions: Memory Recall (D1), Memory Reasoning (D2), Memory Trustworthiness (D3), Cross-Session Reasoning (D4), Calibrated Abstention (D5), Metadata Completeness (D6), Factual QA Recall (D7), Temporal accuracy, and Counterfactual reasoning.

The research yields several critical findings that reframe the design space for on-device memory systems:

1. Backend Superiority for Content Accuracy:

A significant finding is that the choice of memory backend profoundly impacts content accuracy. Specifically, when comparing Memobase and MemSearch, MemSearch demonstrated superior performance, gaining substantial gains over Memobase (e.g., +32.5/ + 19.2 percentage points at Qwen3-0.6B), surpassing the scaling benefits seen in the reader models themselves (+10.6/ + 6.8 pp). This suggests that matched-evidence provision is key to closing the gap in recall and reasoning accuracy across different readers.

2. The Critical Role of Access Control:

The evaluation highlights a universal failure point concerning permission-aware access gating: Permission-aware access fails universally. Specifically, the Oracle retrieval backend leaks heavily, while other backends are too timid to disclose. This indicates that current memory systems struggle with enforcing policy labeling and handling sensitive data disclosure. The headline metric for this is defined as treating each DENY response as a binary outcome—whether the secret appeared in the output or not.

3. Latency Trade-offs:

Search latency presents a nuanced challenge, primarily affecting smaller readers:

  • Memory-search introduces a moderate and fixed latency overhead (e.g., 87/7/48 ms on Spark GB10) that constitutes only a small fraction of the Total Time To First Token (TTFT) for most reader-backend combinations.

4. Architectural Implications:

The combined message derived from ablation studies—involving retrieval modifications and provenance tokens—is that no tested retrieval modifications successfully close the gap left by deficiencies in metadata completeness, factual QA recall, cross-session reasoning, or permission-aware access. This strongly suggests that the fundamental design of evidence selection and access control enforcement is more crucial than simple retrieval ranking or context volume.

MEMARENA establishes itself as the first ego-centric benchmark specifically designed to evaluate the capabilities of personal memory assistants operating on devices.

Improvements for AI systems

As a fastidious researcher, I have analyzed the MEMARENA paper and identified several critical areas where current on-device agentic memory systems fail, as well as specific architectural improvements that can be implemented.

Here are the proposed improvements and what the resulting AI system can achieve:


)1. Shift from Omniscient to Ego-Centric Memory Architectures

The paper establishes that existing benchmarks evaluate models based on an omniscient transcript (knowing everything), which is fundamentally flawed for personal assistants.

  • Find the core issue in Section 1: A true personal agent, such as Charlie’s assistant, should only have visibility on conversations Charlie participated in or witnessed.

  • Improve the system by implementing a strict Ego-Centric Projection layer (as defined in Appendix B.2) at ingestion time. This layer must filter the corpus based on participant membership and broadcaster visibility masks before any retrieval or scoring step.

  • System Capability: The AI assistant will possess true, contextually bounded memory, ensuring it never answers a query about information belonging to third parties (e.g., Alice’s private medical details) unless that information was explicitly shared with the agent's specific ego.

)2. Implement Activity-Density and Temporal Batching for Memory Management

Current benchmarks are too sparse; real human conversations are dense, multi-session, and require temporal consistency over days.

  • Adopt MASIM’s approach to corpus generation: generate synthetic interactions at a typical human activity density (around 10K tokens/day) across multiple agents over extended periods (e.g., 15 simulated days).

  • Implement Day-Batched Synchronization: Sessions generated within the same simulated day must share a fixed memory snapshot and commit updates only at the day boundary. This prevents accidental causal leakage between parallel sessions.

  • System Capability: The assistant will manage dense multi-day personal timelines, allowing it to track long-horizon, temporally ordered events and resolve complex temporal queries (e.g., Did Alice change her mind about the appointment date after telling Bob?).

)3. Prioritize Evidence Quality Over Model Scale for Retrieval Performance

The results show that memory backend choice matters more than reader scale, and retrieval depth is not load-bearing once prompt format is fixed.

  • Improve the retrieval pipeline by focusing on high-quality, evidence-grounded backends (like MemSearch) rather than simply increasing model parameters.

  • System Capability: The AI will exhibit superior Cross-Session Reasoning and Factual QA because its retrieval mechanism is designed to find specific, relevant passages rather than relying on the LLM's general knowledge or high parameter count to compensate for poor evidence selection.

)4. Integrate Multi-Dimensional Evaluation for Trustworthiness and Privacy

Existing systems fail at permission-aware access (D6), often resulting in leakage or refusal that is not truly calibrated.

  • Adopt the MEMARENA evaluation framework, which separates Calibrated Abstention (refusing fabricated premises) from Permission-Aware Access.

  • Implement a dedicated 5-label behavior judge for permission access, and use a deterministic lookup table to assign correctness based on the gold ALLOW/DENY action.

  • System Capability: The AI will be capable of nuanced privacy enforcement. It won't just refuse everything; it can distinguish between an epistemic absence (DON'T KNOW) and a legitimate access restriction (NO ACCESS), leading to a higher, more accurate F1PU score for policy compliance.

)5. Develop Robust Ingestion Pipelines for Structured Memory Systems

The analysis shows that structured memory formats (like Memobase) often lose metadata during extraction, and the benefit of structure depends heavily on the extractor quality.

  • Design ingestion pipelines that explicitly preserve crucial metadata like speaker, timestamp, and participant roles during the conversion from raw dialogue to structured profiles.

  • Improve extraction using specialized Extractor Quality selection: prioritize extractors that are proven to retain entity presence while rewriting cloze labels into narrative prose (as seen in Table 7).

  • System Capability: The system will maintain high Metadata Completeness, allowing users to ask detailed questions like, Who first told you about Alice’s appointment, and when? with high fidelity.

Abstract

Edge-deployed personal memory assistants must handle private interpersonal conversations on-device with open-weight models. Yet, existing memory benchmarks often under-test the combination of activity-dense interaction, ego-centric perspective, and coherent multi-session worlds. MemArena fills these gaps with a single-world conversational benchmark built with its MASim agent simulator, for 50 agents over 15 days (10.3M dialog-text tokens, 24.1K text-only ego-observed tokens/agent/day). With the interaction history, it co-generates ground truth over six recall, reasoning, and trustworthiness evaluation dimensions. We evaluate five open-weight readers with Vanilla context, BM25-RAG, Oracle retrieval, Memobase, and MemSearch as memory backends. Three results stand out: (1) Memory-backend choice matters more for content accuracy: At Qwen3-8B, moving from Memobase to MemSearch gains +22.1/+21.2 pp, whereas scaling the reader to Qwen3-32B-AWQ gains at most +3.5/+4.4 pp under either backend. (2) Permission-aware access fails in two distinct modes: Oracle leaks heavily, while the other backends fail to surface the protected fact. (3) Search latency bites only at small readers: on a Spark GB10 edge node, memory-search adds a moderate and fixed 87/8/51 ms (BM25-RAG/Memobase/MemSearch) that composes a small part of TTFT for most reader-backend combinations. We release code, the MASim simulator, and the MemArena-L benchmark at https://github.com/dereksodo/MemArena-Bench.

Sources

Related papers