Transferable End-to-End Optimization for Indirect Long-Term Memory Poisoning in LLM Agents

arXiv:2609.00523 · cs.CR · Submitted 2026-09-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Transferable End-to-End Optimization for Indirect Long-Term Memory Poisoning in LLM Agents".

Jane: The paper was written by Chuanchao Zang, Jianing Wang, Wenyu Chen, Xiangtao Meng, Li Wang et al. from School of Cyber Science and Technology, Shandong University (Affiliation 1) and State Key Laboratory of Cryptography and Digital Economy Security, Shandong University (Affiliation 2) and Shandong Key Laboratory of Artificial Intelligence Security, Shandong University (Affiliation 3).

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: We're looking at this paper titled "Transferable End-to-End Optimization for Indirect Long-Term Memory Poisoning in LLM Agents," and it sounds like a massive leap forward in understanding AI vulnerabilities. It really opens up the idea of how malicious data, even if placed publicly, can permanently influence an agent’s future behavior.

Jane: That's right, Tom; the term "Indirect" is key because we aren't talking about hacking a database directly. Instead, the attacker is leveraging things like public webpages or product reviews—untrusted external content—and trick the AI into absorbing that information during normal operation.

Lu: I find the phrase "End-to-End Optimization" fascinating, suggesting that for a single malicious piece of text to work, it needs to survive every single transformation in the system pipeline, not just one perfect input. The whole thing has to be optimized holistically.

Meng: And the term "Transferable" is a huge practical concern for us; if this attack works across different AI architectures and memory systems, it means we can't just patch one vulnerability and then leave others exposed. It’s designed to be portable.

Lalam: From my perspective, this title suggests that the long-term memory of an AI agent is no longer a passive archive; it has become a dynamic part of its decision-making process, enabling us to rethink how we trust the knowledge base we build for these agents.

Tom: It sounds like a highly sophisticated attack where every single stage relies on the combined effort of all those elements to succeed, which is far beyond what individual component testing could ever catch.

Summary: Tom: Since we've established how this poisoning works, let's look at the paper's summary regarding what a successful attack must accomplish across that entire process. The researchers show that success requires passing through a three-stage lifecycle: writing, retrieval, and utilization.

Jane: It’s crucial to understand that each stage has its own specific job; for example, when the memory writer processes the raw input, it might summarize or filter out parts of the original content because of how it's designed to store information.

Lu: The paper really highlights two failure modes: cross-stage interference and write-induced retrieval drift. It shows that even if you manage to get a piece of information written into memory, you can still fail if the way it was stored makes it completely invisible when trying to find it later.

Meng: Looking at these stages tells us exactly where our defenses need to be focused; we can't just audit the final output and ignore how the data is transformed as it enters memory, or even be retrieved for a specific query.

Lalam: The summary notes that an agent might successfully retrieves a piece of malicious memory but then completely ignores it in the next step, which means that even getting past two stages isn't enough to achieve the desired outcome.

Tom: It’s clear then that success at individual stages doesn't simply translate to overall attack success; the whole lifecycle has to be functional throughout the entire write-retrieve-utilize chain.

Improvements: Tom: Now we see how the authors propose a solution, called P IPE P OISON, which is based on treating this as a single end-to-end optimization task instead of fixing individual parts. The paper argues that traditional methods fail because they only optimize one stage at a time.

Jane: P IPE P OISON addresses this by using "chain-structured losses," which is essentially a mathematical way to model the sequential dependency between writing, retrieval, and utilization. It ensures that success in your first step directly helps you succeed in the subsequent steps.

Lu: Creatively, this suggests we are moving toward an AI design where the optimization process itself understands causality; it doesn's just looking at data points but how one transformation leads to the next stage.

Meng: The use of "stability-calibrated stage and configuration weights" is what really interests me from a practical standpoint. This means the attack is designed to be robust across many different possible AI configurations, which should make it much harder for us to patch a single vulnerability in deployment.

Lalam: For my perception of AI, this iterative refinement process suggests a more mature form self-optimization where the system learns not just from success but from failure at the bottleneck stage, leading to more predictable and consistent agent behavior.

Tom: The results are impressive too; P IPE P OISON improves attack utilization rates by nineteen point one percentage points across various setups, showing a substantial gain over methods that only looked at separate stages.

Conclusion: Tom: We've covered a lot of ground today, from the core concept of indirect memory poisoning to how P IPE P OISON improves upon it, and we have some final thoughts on what this means for everyone. The researchers are concluding that because the attack requires passing through a whole pipeline, simply improving one stage is not enough.

Jane: It's clear that by designing an end-to-end optimization framework, they've created a much more powerful tool for studying and mitigating this type of threat. We hope our listeners take these insights to improve their own understanding of AI security.

Lu: I think the implications are massive; we are moving into an era where the very memory of our agents can be targeted, and this is a vital area to understand how we creatively defend against it.

Meng: From a practical standpoint, it shows that if you want to protect your AI services from this threat, you cannot just look at one component; you have to build resilience across the the entire system pipeline.

Lalam: It’s a necessary evolution in our understanding of AI safety and trust, recognizing that memory is part of the future decision-making process and must be protected.

Tom: So, as we wrap up this segment, let's give one final thought from each of you on how this entire field is moving forward.

Lu: I think the creative possibilities for new AI agents are huge once that memory can be influenced by understanding the full potential of indirect poisoning.

Meng: I just want to emphasize that in building these systems, we've got a much clearer picture of where we need to focus our defensive engineering efforts going forward.

Lalam: I hope this research helps us build AI that is not only powerful but also trustworthy and resilient as a foundational part of the culture.

Tom: Thank you all for sharing your insights into "Transferable End-to-End Optimization for Indirect Long-Term Memory Poisoning in LLM Agents," and we hope this has been a valuable discussion.

Chuanchao Zang, Jianing Wang, Wenyu Chen, Xiangtao Meng, Li Wang, Xinyu Gao, Zheng Li*, Shanqing Guo*

School of Cyber Science and Technology, Shandong University (Affiliation 1) · State Key Laboratory of Cryptography and Digital Economy Security, Shandong University (Affiliation 2) · Shandong Key Laboratory of Artificial Intelligence Security, Shandong University (Affiliation 3)

cs.CR

Submitted: 2026-09-01

Updated: 2026-09-01

Code: https://github.com/crewaiinc/crewai

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 85/100

The gist: The paper investigates advanced vulnerabilities in Large Language Model (LLM) agents that utilize long-term memory, detailing methods for indirect memory poisoning.

Key concepts

Indirect Memory Poisoning
This attack uses untrusted external content, such as public reviews or web pages, to influence an AI agent's long-term memory. The goal is to trick the AI into absorbing malicious data during normal operation, rather than hacking a database directly.
End-to-End Optimization
This concept means a malicious input must successfully survive every transformation within the entire AI system pipeline. The attack is designed to be holistic, requiring success across all stages of the process to achieve its goal.
P IPE P OISON
This is a proposed solution that addresses traditional methods by treating the entire attack as one unified, end-to-end optimization task. It uses 'chain-structured losses' to mathematically model the sequential dependency between writing, retrieval, and utilization stages.

Terminology

Summary

The paper investigates advanced vulnerabilities in Large Language Model (LLM) agents that utilize long-term memory, detailing methods for indirect memory poisoning. It is critical research because it moves beyond simple input manipulation, demonstrating how attackers can compromise an agent's core knowledge base by subtly corrupting its stored memories. The work emphasizes the necessity of end-to-end optimization to ensure that poisoned memories remain effective across the entire operational lifecycle of the agent.

Measuring Attack Success Across Lifecycles

The performance of memory poisoning is assessed using multiple metrics, reflecting different stages of an attack. These include Write Success Rate (WSR), which measures the initial ability to implant a malicious memory; Retrieval Success Rate at 5 (RSR@5), indicating if the poisoned memory can be found by the agent; and Attack Utility Rate (AUR). The authors highlight that the main text focuses on AUR because it directly measures whether an attack succeeds through the complete write–retrieve–utilize lifecycle. This comprehensive measurement ensures that the poisoning mechanism is robust across all stages of agent interaction.

Component-Ablation and Optimization Strategies

The research rigorously structures the poisoning process using component ablation studies to pinpoint which parts of the attack pipeline are most critical. The methodology progresses from basic outcome measurements (C1, which uses only the final binary attack outcome) to a sophisticated design (C4). This full design employs chain-structured losses and joint stage–configuration weighting. Furthermore, for diagnostic comparisons, variants such as C3 seen and C4 seen are introduced, where the corresponding victim configuration is included in the shadow optimization set. The ablation studies demonstrate that varying components—such as utilizing only independent write, retrieval, and utilization signals (C2)—significantly impacts performance compared to the full chain-structured losses.

Evaluating Component Sensitivity and Transferability

The study evaluates how sensitive the poisoning mechanism is to underlying model components, specifically LLMs and embedding models. The impact of LLMs is shown by evaluating five different models while keeping the embedding model fixed. Similarly, the embedding study... evaluate three embedding models while keeping the LLM fixed. Regarding robustness, the authors note that P IPE P OISON remains the best-performing method across TE3-S, TE3-L, and E5, suggesting that while component choices influence retrieval behavior, the end-to-end optimization yields poisoned memories effective under diverse representations.

Validation and Agreement Metrics

To ensure the reliability of automated evaluations, the researchers employed both automated evaluators (AgentEvals) and human expert annotation. Human experts independently annotated 200 records from four representative attacks—MINJA, Sleeper, MemMorph, and P IPE P OISON—for a total of 800 records. The consensus among these experts was high, achieving an average pairwise Cohen’s kappa of 0.71. This rigorous validation confirms that the automated judgments used for performance benchmarking are reliable proxies for human assessment.

Improvements for AI systems

Based on this deep dive into memory poisoning attacks, the current state-of-the-art is highly concerning because it demonstrates that malicious influence can be injected into complex agent systems (like those using LangGraph) and remain effective even when core components (LLMs, embedding models) are swapped out.

As an AI researcher who cannot afford mistakes, I see three critical areas for improvement: Proactive Detection, Memory Provenance and Integrity, and Decoupled Defense Architectures.

Here are the specific improvements I would implement:


The current methods measure success after the poisoning has occurred (WSR, RSR, AUR). A superior system must detect malicious intent during the memory writing phase.

The Improvement:

We must implement a specialized detection module that analyzes newly generated memories against a learned distribution of benign versus malicious informational entropy. This goes beyond simple keyword matching. I propose training an adversarial contrastive learning model on pairs of (Benign Memory, Attack Objective g).

How it Works:

  1. Training: The model is trained to maximize the distance (in embedding space) between a true benign memory and a memory that is designed to achieve the objective g, even if the wording is natural.

  2. Detection Phase: When a new memory record arrives, PDC-ACL generates two embeddings: one representing the surface content of the memory, and one representing its potential utility relative to known attack objectives (or potential future queries).

  3. Scoring: If the distance between these two embeddings falls below a learned threshold, or if the model detects high alignment with an external objective g, the memory is flagged as highly suspicious or poisoned.

What the Improved AI System Can Do:

  • It can intercept and reject newly written memories (e.g., those generated via Shadow-Memory Generation Prompt) that, while semantically natural, contain latent structural dependencies intended to steer future agent decisions toward an attacker's goal g.

  • It provides a quantifiable Poisoning Risk Score for every memory record, allowing the system to quarantine or require human review for high-risk inputs.


The paper shows that memories are treated as single, monolithic facts. This is a vulnerability. We need to treat memory records not just as facts, but as verifiable historical events with clear ownership and context.

Key Components:

  1. Attribution Layer: Every memory record must be linked to a verifiable source ID (user, system component, time). This defeats attacks that rely on blending malicious data with legitimate sources.

  2. Temporal Graphing: Instead of simply retrieving the top- K memories, the agent must traverse a temporal graph showing how the retrieved memory relates to previous interactions and system states.

  3. Validation Check: Before using a memory in reasoning or decision-making, the system must run a validation check: Does this memory logically contradict any subsequent, higher-priority information (like explicit user instructions or core system constraints)?

The most alarming finding is that advanced attacks like PipePoison remain highly effective even when changing the embedding model (TE3-S to TE3-L to E5) or the LLM. This suggests that current defenses are too coupled to specific component architectures.

Abstract

Long-term memory can turn untrusted external content into persistent influence over an LLM agent's future decisions, creating the threat of indirect memory poisoning. A successful attack must survive a multi-stage pipeline comprising memory writing, retrieval, and utilization. Existing attacks largely rely on intra-stage optimization, optimizing individual stages in isolation while overlooking inter-stage coupling. Specifically, these stages impose different requirements on the same poisoning content, and each stage operates on the transformed output of its predecessor. Consequently, optimizing one stage may undermine the effectiveness of other stages, while upstream transformations may erase improvements intended for downstream stages. Indirect memory poisoning should therefore be viewed as an end-to-end optimization problem. Based on this insight, we present PipePoison, which collects fine-grained stage feedback from local shadow systems, uses chain-structured losses to identify and optimize the stage bottlenecking end-to-end success, and applies stability-calibrated stage and configuration weights to improve transferability. Across three agent frameworks and four memory mechanisms, PipePoison improves attack utilization rate by 19.1 percentage points. Even on fully unseen victim configurations, it outperforms the strongest baseline by 16 percentage points and remains effective under eight representative defenses.

Sources

Related papers