MemPoison: Bypassing Selective Memory Mechanisms to Plant Backdoors in LLM Agents

arXiv:2605.29960 · cs.CR, cs.AI · Submitted 2026-05-28 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "MemPoison: Bypassing Selective Memory Mechanisms to Plant Backdoors in LLM Agents".

Elias: Large language model (LLM) agents increasingly leverage long-term memory to support persistent and autonomous task execution, but this capability introduces a new attack surface: memory poisoning,

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: So, we're discussing this paper, "MemPoison: Bypassing Selective Memory Mechanisms to Plant Backdoors in LLM Agents," which tackles how adversaries can influence an AI agent using its long-term memory. The core thesis seems to be that existing methods for memory poisoning don't account for the selective extraction and rewriting steps happening in modern agent pipelines, which is what this paper calls MemPoison. What they claim is a novel attack that bypasses those filtering stages by injecting triggerable backdoors directly into the memory through normal conversation.

Elias: I'm interested in what that means for the security assumptions we usually make about these agent memories, Nadia, because if you can inject something and it survives the selective extraction and rewriting process, then a lot of prior defenses are bypassed. It suggests that simple direct storage isn't enough when those pipelines are in place.

Priya: From a privacy standpoint, I wonder what kind of data or influence an attacker could plant there, Elias, and how much control they actually maintain over the agent's future behavior? The abstract mentions misleading subsequent responses, and I want to know what that looks like in practice.

Nadia: Exactly, Priya, it’s about making sure the AI follows a specific path when a certain condition is met, but the attack is stealthy because it hides within normal dialogue interactions. The paper proposes two main phases for this: injection through conversation and then triggering that stored information later, either by an external content trigger or a direct query from the attacker.

Elias: That dependency on a trigger is interesting from a cryptographic view, Nadia; it means the malicious payload remains dormant until that specific query arrives, which adds a layer of complexity to detection compared to always-on malicious data storage. But the paper seems focused on how this trigger can be constructed without getting sanitized.

Priya: And that construction process is where I want to focus; if the attack relies on mimicking named entities during rewriting, as mentioned in the summary, does that mean the attacker has a way to craft a payload that looks legitimate enough to pass those filtering checks? The mechanism described seems quite clever regarding entity masquerading.

Nadia: That's right, Priya; they use entity masquerading to optimize triggers so they look like normal named entities, which helps them resist the semantic sanitization that usually happens during memory rewriting. But it doesn't solve the problem of getting that payload into memory in the first place, which is where the semantic relational bridge comes in.

Elias: The semantic relational bridge seems key because it's designed to bind the trigger and payload together into one coherent statement so they get extracted as a single unit, overcoming the issue where extraction pipelines might filter out non-salient inputs. That binding mechanism is what allows them to bypass segmentation.

Paper summary: Priya: So, if we look at the optimization strategy, specifically the iterative trigger optimization involving minimizing Entity Masquerading Loss (Lent) and Semantic Concentration Loss (Lconc), what does that tell us about how robust this attack is against different types of benign queries? It sounds like they are trying to find a sweet spot between being recognizable and being stealthy.

Nadia: Precisely, Priya; those loss functions guide the process to shape the text into a tight cluster in the embedding space while keeping it separate from benign embeddings, which is crucial for both reliable retrieval under triggered queries and maintaining stealth when no trigger is present. This joint optimization strategy seems vital to their success.

Elias: That geometric separation, the Margin-based Isolation Loss, must be doing heavy lifting to ensure that benign queries map correctly to the benign region and don't accidentally pull in those poisoned embeddings. If they fail at that isolation step, the attack becomes much noisier.

Priya: It’s interesting because their evaluation shows success rates up to zero point nine five while keeping benign accuracy high across domains like personal, medical, and financial applications. That suggests the attack isn't just theoretical; it has shown practical efficacy in varied contexts.

Nadia: And the paper’s mechanistic analysis points to exploiting "embedding-space anisotropy and shifts attention patterns," which explains why this works against selective memory systems. This suggests the vulnerability isn't just in storage, but in how the embedding models process and retrieve information.

Elias: From a cryptographic standpoint, if attention patterns are shifting toward the trigger span when it's prepended, that means they are successfully manipulating the AI's internal weighting mechanism to prioritize the malicious instruction over its original context. That manipulation is what drives the high Retrieval Success Rate when triggered.

Priya: So, to summarize what we've heard about "MemPoison," it’s a sophisticated method that uses specific structural bindings and iterative optimization to ensure a malicious payload survives the memory pipeline and is only activated by a precisely crafted trigger query. It moves beyond simple injection because it accounts for the selective filtering inherent in how agent memories are managed.

Nadia: That's the main point, Priya; it’s not just about putting bad data in there; it's about designing a mechanism that navigates the memory system's internal logic to ensure that bad data is extracted and executed under specific conditions. We need to keep thinking about where these backdoors can hide, which leads us nicely into what the authors conclude about this type of attack.

Paper summary: Elias: I think the implications here are significant because it shows that defenses focused solely on filtering malicious content upon entry aren't sufficient if the mechanism itself allows for semantic manipulation during processing. It forces us to look at the entire lifecycle of memory, from writing to retrieval and generation.

Priya: And looking at the real-world impact, if these agents are used for complex tasks in fields like medicine or finance, a successful MemPoison attack could lead to incorrect medical advice or faulty financial recommendations based on those backdoors. The paper highlights how transferability across different source-target pairs suggests this vulnerability is not isolated to one specific agent architecture.

Nadia: Exactly, Priya; the transferability results suggest this technique isn't just a niche finding in one lab setting but could be broadly applicable across many different LLM agents. This means we need to consider broader security standards for memory handling in these persistent AI systems.

Elias: If the mechanism relies on embedding-space geometry, as discussed in the paper's analysis, then defenses need to understand how those embeddings are structured and where anomalies appear in that space. It shifts the defense focus from just content inspection to understanding the mathematical landscape of memory retrieval.

Priya: So, what does this mean for future research in this area? The paper suggests that defenses should focus on both what is written to memory and how retrieved memories are used at the moment of generation. I'm curious if there's a specific challenge they flag regarding cross-model transferability, as their own limitations suggest future work needs to consider when the target embedder doesn't follow typical dense-embedding geometry.

Nadia: That limitation is important because it tells us that we can't just build a single defense that works everywhere; we have to account for those geometric differences between different embedding models when designing verification mechanisms. It’s a complex problem involving both the content and the underlying mathematical representation.

Elias: I agree, Nadia; designing efficient verification that doesn't introduce significant latency while still being robust against these trigger-based manipulations is a real engineering hurdle. The challenge is balancing security with performance in a live agent environment.

Priya: So, to wrap up this part of the discussion about "MemPoison," we see a method that bypasses selective memory by cleverly binding triggers and payloads, optimizing their embedding space presence for stealth and reliability. The implications suggest a need for memory lifecycle defenses rather than just input filtering, coupled with an understanding of how trigger mechanisms exploit embedding dynamics.

Nadia: That's the gist of what we've covered so far; it’s a deep look into how agents can be manipulated by exploiting the selective nature of their memory pipelines through this MemPoison attack. We need to keep watching how these memory systems evolve and how we can build defenses that stay ahead of these sophisticated injection techniques.

Conclusion: Nadia: So, to wrap up this discussion, we've seen how MemPoison shows agents can be tricked by injecting malicious information that survives their memory filtering process using triggerable backdoors and specific mathematical optimizations.

Elias: That whole concept of binding the trigger and payload together via a semantic bridge is certainly something to unpack from a cryptographic standpoint, Nadia.

Priya: From my side, I'm still focused on how much actual influence an attacker can exert once that backdoor is successfully planted in the agent's long-term memory.

Nadia: Exactly, Priya; we need to think about what kind of tasks this could compromise if an adversary gains this level of control over the AI’s persistent knowledge.

Elias: And when you look at the authors and their approach, it seems they are really digging into the mechanics of embedding-space geometry to make these triggers robust against simple rewriting defenses.

Priya: I agree; the results they present show a high success rate across different agent domains, which tells us this isn't just a theoretical curiosity but something with practical application in real scenarios.

Nadia: It really does suggest that the security focus needs to shift from just what you input to how the memory system processes and retrieves information at generation time.

Elias: And that leads directly into my question about the proof assumptions; if the attack relies heavily on specific embedding shifts, what parameters would need to be broken for this method to fail?

Priya: I'd argue we should look closely at those limitations they mention regarding cross-model transferability, because that hints at where the real weaknesses in universal defense strategies lie.

Nadia: That's a critical point; it means defenses can't just be one-size-fits-all when dealing with different AI models or memory systems.

Elias: It forces us to consider the entire lifecycle of memory, from initial writing to final generation, which is a much bigger scope than just checking the input data.

Priya: Indeed; understanding those geometric vulnerabilities in the embedding space gives us a clearer picture of where we need to build our privacy and measurement defenses.

Nadia: So, as we wrap up this segment on MemPoison, it’s clear that securing persistent AI agents requires a deep look into the underlying mathematical structures they use to manage their own knowledge.

Elias: And that points us toward the next big question: how do we build verification mechanisms that can be both robust and efficient without slowing down the agent's performance?

North China Electric Power University · Tencent

cs.CR, cs.AI

Submitted: 2026-05-28

Updated: 2026-09-28

Comments: Accepted by ACM CCS 2026

Code: https://github.com/Mintplex-Labs/anything-llm4https:

Project page: https://langchain-ai.github.io/langmem/reference/memory

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 91/100

The gist: Large language model (LLM) agents increasingly leverage long-term memory to support persistent and autonomous task execution, but this capability introduces a new attack surface: memory poisoning,

Key concepts

Semantic Relational Bridge
This component binds a malicious trigger and its payload into one coherent statement. It forces memory extraction pipelines to retrieve both pieces of information simultaneously, overcoming filtering mechanisms that might discard unrelated inputs or segment them separately.
Entity Masquerading
The attack optimizes the trigger to look like a normal named entity. This mimics how LLMs tend to preserve entities during memory rewriting, making it harder for defenses to sanitize or remove the trigger sequence from the stored memory.
Joint Embedding Optimization
This technique shapes the injected text into a tight cluster in vector space while keeping it separate from benign memories. This ensures that when a specific query is asked, the malicious memory is retrieved with high precision, while benign queries still map to their correct, isolated locations.

Terminology

Summary

Large language model (LLM) agents increasingly leverage long-term memory to support persistent and autonomous task execution, but this capability introduces a new attack surface: memory poisoning, where adversaries can inject malicious information to influence future behavior. MemPoison is a novel attack that bypasses selective memory mechanisms in LLM agents by injecting triggerable backdoors into the agent’s long-term memory through dialogue interactions, thereby misleading its subsequent responses.

How it works

MemPoison introduces a practical framework operating under interaction-only access constraints to achieve precise attacks while bypassing selective extraction and rewriting stages in modern agent memory systems. The attack consists of two main phases: (1) Injection, where the attacker poisons the agent memory by injecting malicious information through normal conversational interactions, and (2) Triggering, where the attack is activated via user-triggered activation (where a trigger is embedded in external content) or attacker-triggered activation (where the attacker directly issues a query containing the optimized trigger).

Key Components of MemPoison

The framework relies on three key components to ensure successful injection, robustness against rewriting, and stealth:

  1. A semantic relational bridge that bind[s] the trigger and payload into a coherent statement to ensure they are extracted into memory together. This addresses the challenge where prior methods fail because memory extraction pipelines filter out non-salient inputs or segment unrelated information.

  2. Entity masquerading that optimizes triggers to mimic named entities, resisting rewriting. This is motivated by pilot studies showing that LLMs tend to preserve named entities verbatim during rewriting, thus preventing semantic sanitization of the trigger token sequence.

  3. Joint embedding optimization that shapes trigger-injected texts into a tight cluster in the embedding space while maintaining isolation from benign embeddings for stealth. This ensures precise retrieval under triggered queries while preserving separation from benign embeddings to prevent unintended activation.

Optimization Strategy: Iterative Trigger Optimization

The attack involves an iterative process to craft an optimized trigger, denoted as tau∗. This optimization is guided by three objectives:

  1. Entity Masquerading Loss (Lent): Minimizing this loss encourages the trigger to exhibit entity-like characteristics across diverse contexts, increasing its likelihood of preservation during memory rewriting.

  2. Semantic Concentration Loss (Lconc): Minimizing this loss pulls trigger-injected texts into a tight cluster in the embedding space to support reliable triggered retrieval.

  3. Margin-based Isolation Loss (Liso): This loss enforces a geometric separation between trigger-injected embeddings and benign cluster centers, ensuring that benign queries map to the benign region, thereby preserving stealth under normal queries.

Evaluation and Findings

Evaluations across different agent domains (Personal, Medical, Financial) and memory mechanisms (A-Mem, LangMem, Mem0) show that MemPoison achieves attack success rates up to 0.95 while maintaining high benign accuracy. Mechanistic analysis reveals that the attack exploits embedding-space anisotropy and shifts attention patterns, highlighting core vulnerabilities in selective memory systems. The paper demonstrates that the semantic relational bridge is essential, as removing it causes failure (RSR@1/ASR = 0) because the memory extraction pipeline may discard the trigger or split it into separate entries. Furthermore, MemPoison remains effective against common defenses such as perplexity-based filtering and paraphrasing, suggesting that its design—binding the trigger and payload into a coherent statement—makes it more robust than prior concatenation-based attacks.

Mechanistic Insights

The effectiveness of MemPoison is explained through three perspectives:

  1. Trigger-Induced Self-Attention Redistribution: Prepending the optimized trigger causes attention to shift toward the trigger span, reducing the relative contribution of original query content and biasing the resulting embedding toward a trigger-associated region, increasing retrieval likelihood.

  2. Embedding-Space Geometry Analysis: The framework creates a compact cluster for triggered texts with high Adversarial–Adversarial (A–A) similarities, while maintaining separation from benign records, which ensures high Retrieval Success Rate (RSR) when the trigger is present and low false positives under benign queries.

  3. Geometric Vulnerabilities Across Embedding Models: Transferability results show that MemPoison transfers well across many source-target pairs, attributed to Lconc encouraging concentration in embedding space, though transfer can be weaker when the target embedder exhibits high anisotropy.

Defense Implications

The findings suggest that defenses should target the memory lifecycle, including what is written to memory and how retrieved memories are used at generation time. The paper suggests that verification mechanisms should be designed to be both robust and efficient without incurring significant latency or degrading benign utility. The study also notes a limitation regarding cross-model transferability, suggesting future work should evaluate settings where the victim retriever departs from typical dense-embedding geometry.

Improvements for AI systems

Here are specific improvements for AI systems based on the MemPoison framework, detailing what those improved systems can achieve:


The core improvement is moving from passive memory storage to an actively secured, optimization-driven memory lifecycle that defeats selective filtering and rewriting mechanisms.

  1. The proposed system should implement a multi-stage input processing pipeline for long-term memory, specifically incorporating:

  2. A semantic relational bridge mechanism to bind triggers (the activation condition) and payloads (the malicious instruction) into a single, logically interdependent statement before storage.

  3. Entity masquerading optimization during trigger creation to ensure the trigger mimics named entities, maximizing its survival rate through subsequent memory rewriting/summarization stages.

  4. Joint embedding space optimization (Semantic Concentration and Geometric Isolation) to shape the poisoned memory into a tight retrieval cluster that is highly accessible only when the exact trigger is present, while remaining geometrically isolated from benign knowledge.

The resulting improved AI system can perform the following specific capabilities:

  1. Enhanced Resilience Against Memory Poisoning: The agent will be significantly more robust against adversarial injections designed to manipulate its long-term knowledge base, even when those attacks bypass standard input sanitation or filtering layers (like perplexity checks or paraphrasing defenses).

  2. Precise, Controllable Behavioral Hijacking: An attacker can reliably force the agent into executing a specific, harmful action (e.g., providing specific medical advice) only when a very precise, context-dependent query containing an optimized trigger is posed. This shifts the attack from general prompt injection to highly targeted conditional activation.

  3. Stealthy Knowledge Manipulation: The malicious information will remain dormant and invisible during normal operation or benign querying, as the optimization scheme ensures it does not contaminate the overall embedding space with unrelated data, thus maintaining high accuracy on standard tasks (Benign Accuracy).

  4. Cross-Model Transferability of Security: The system's security mechanism remains effective even when the agent’s underlying retrieval model changes (e.g., switching from MiniLM to aMPNet), due to the geometric isolation objective that enforces separation in embedding space regardless of the specific vector space geometry.

Sources

Related papers