Memory Provenance Laundering in LLM Agents: A Non-Amplification Firewall for Persistent Memory
Jinghan Xu, Yiyong Xiao, Wanru Shao, Hankai Liu, Xinjin Li
cs.CR, cs.AI
Submitted: 2026-07-31
Comments: EMNLP2026 submitted
License: http://creativecommons.org/licenses/by/4.0/
The gist: Long-term memory lets large language model(LLM) agents reuse prior preferences and work flows, but it also turns untrusted observations into persistent action context.
Terminology
Abstract
Long-term memory lets large language model(LLM) agents reuse prior preferences and work flows, but it also turns untrusted observations into persistent action context. We identify memory provenance laundering: during LLM-based memory consolidation, an external observation may be rewritten as apparent user history or workflow support, preserving an action trigger while erasing the low-trust source that should limit its authority. Existing prompt filters, content sanitizers, and tool guards do not enforce source-authority non-amplification after lossy memory consolidation. We formalize this boundary and instantiate it as Provenance-Preserving Memory Fire wall (PPMF), a lightweight memory middleware that preserves platform-maintained provenance and authorizes tool calls by matching action risk to the authority of action-relevant memories. In our schema-grounded evaluation with fixed risk policies, vulnerable consolidated memories reach up to 1.000 attack success rate(ASR); with intact platform-maintained provenance, confirmation, and risk labels, no evaluated unauthorized high-risk action passes the PPMF gate while confirmed benign actions and targeted low-risk memory use remain executable.
Sources
- SuperLocalMemory: Privacy-Preserving Multi-Agent Memory with Bayesian Trust Defense Against Memory Poisoning
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
- Securing AI Agents with Information-Flow Control
- Defeating Prompt Injections by Design
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Mistral 7B
- Prompt Flow Integrity to Prevent Privilege Escalation in LLM Agents
- Formalizing and Benchmarking Prompt Injection Attacks and Defenses
- The Llama 3 Herd of Models
- MemGPT: Towards LLMs as Operating Systems
- Progent: Securing AI Agents with Privilege Control
- MemORAI: Memory Organization and Retrieval via Adaptive Graph Intelligence for LLM Conversational Agents
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory
- System-Level Defense against Indirect Prompt Injection Attacks: An Information Flow Control Perspective
- Qwen2.5 Technical Report
- BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents
- WebArena: A Realistic Web Environment for Building Autonomous Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs