MemAudit: Auditing Long-Term Agent Memory via Hidden User-State Recovery
cs.CL
Submitted: 2026-06-23
Updated: 2026-10-01
Code: https://github.com/sora1998/MemProbe
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- HaluMem: Evaluating Hallucinations in Memory Systems of Agents
- Bootstrapping a User-Centered Task-Oriented Dialogue System
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks
- Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
- Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale
- PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
- PERMA: Benchmarking Personalized Memory Agents via Event-Driven Preference and Realistic Task Environments
- Evaluating Very Long-Term Conversational Memory of LLM Agents
- MemGPT: Towards LLMs as Operating Systems
- Generative Agents: Interactive Simulacra of Human Behavior
- LaMP: When Large Language Models Meet Personalization
- Reflexion: Language Agents with Verbal Reinforcement Learning
- JudgeBench: A Benchmark for Evaluating LLM-based Judges
- From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents
- DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas
- LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
- AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment
- A-MEM: Agentic Memory for LLM Agents
- PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering