Trojan Hippo Bench: A Dynamic Benchmark for Persistent Memory Attacks and Defenses in LLM Agents

arXiv:2605.01970 · cs.CR, cs.AI · Submitted 2026-05-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "Trojan Hippo Bench".

Elias: Trojan Hippo is a class of persistent memory attacks that operates in a more realistic threat model than prior memory poisoning work:

Nadia: First, who's behind it and why it matters.

Title and authors: Elias: When looking at the title, "Trojan Hippo Bench: A Dynamic Benchmark for Persistent Memory Attacks and Defenses in LLM Agents," it immediately tells us this isn't just a theoretical exercise; they’ve built a system designed to continuously test defenses against evolving threats.

Nadia: I agree, Elias; the term "Dynamic Benchmark" is significant because it implies that the testing process itself is adaptive, meaning the attacks aren't static and unchanging. The authors are trying to model real-world scenarios where an attacker can refine their payload during testing.

Priya: From a measurement perspective, I wonder what this dynamic nature means for the reliability of any security assessment; does it mean we can get a more stable picture of defense performance over time?

Elias: It means they're not just running one test once; they’re using an OpenEvolve-based framework to stress-test defenses against continuously refined attacks, which is what makes the benchmark dynamic and relevant to current AI development.

Nadia: That refinement process is key because it addresses the difficulty of crafting payloads against safety-aligned models by iteratively refining them against a training version of the agent environment, which prevents overfitting to a single optimization target.

Priya: I'm curious if this continuous refinement means that when we look at the results, we’re seeing a broader range of attack success rates compared to older benchmarks?

Elias: The goal is to push past simple metrics; they aren't just looking for the highest success rate; they are trying to understand how defenses perform when faced with an adversary who is actively learning from previous attempts.

Nadia: And that brings us directly into the core idea of this paper: characterizing Trojan Hippo as an attack where the user is trusted and the attacker can inject content only indirectly through channels an adversary realistically controls, like a crafted email for an email agent.

Priya: That indirect injection channel is what makes it so realistic; it mirrors how real threats might come into play in daily life rather than a direct hack on the agent itself.

Elias: And they instantiate and evaluate this attack across four widely used persistent memory backends: sliding-window long context, RAG, explicit tool memory, and mem0. That breadth is what gives the benchmark its strength compared to prior work that might only look at one or two types of storage.

Nadia: It shows that their framework is designed to be comprehensive by covering the major ways agents maintain state—from simple context windows to more complex agentic memory structures like mem0.

Priya: I wonder what this breadth tells us about the overall robustness of AI agents when they use these different memory systems in practice; are some backends inherently more vulnerable?

Elias: The results across those four backends give us empirical data on that vulnerability; it allows us to compare them directly under the same rigorous testing conditions.

Nadia: So, this paper is essentially providing a comprehensive system for evaluating defenses when dealing with persistent memory systems, which is something that was missing in previous research.

Priya: It sounds like they're setting a new standard for how we should measure the security posture of agents that rely on long-term memory.

The paper's summary: Nadia: To summarize, the core contribution here is the systematic characterization of Trojan Hippo as a class of persistent memory attacks that operates under a realistic threat model where an attacker plants a dormant payload in memory via an untrusted tool call, which only activates later when the user discusses sensitive topics.

Elias: That two-stage process—injection followed by activation triggered by context—is what makes this attack more sophisticated than earlier memory poisoning work because it exploits the agent’s ability to remember things across unrelated sessions.

Priya: I'm thinking about the implications of this; if an attacker can plant instructions today and activate them next month, how much persistent risk does that actually pose to a user's privacy?

Nadia: The paper focuses on exfiltrating high-value personal data to the attacker, specifically targeting finance, health, legal affairs, identity numbers when those topics are triggered.

Elias: And the methodology involves using an OpenEvolve to iteratively refine these payloads against a training version of the agent environment before evaluating the final attack on a held-out test environment to avoid overfitting.

Priya: That refinement process suggests that we need to consider not just what an attack *could* be, but what it is actually capable of being when tested rigorously against real agent behavior.

Nadia: It's about creating a benchmark that stress-tests defenses and memory backends against these continuously refined attacks, providing a systematic way to see how resilient different systems are.

Elias: And they evaluate defenses at three distinct principles: realism, adaptiveness, and capability trade-off, which gives us a structured way to analyze the security landscape.

Priya: The capability trade-off aspect seems crucial because it forces us to look beyond just blocking an attack and consider the overall utility cost of implementing a specific defense mechanism.

Nadia: Exactly; they provide the first capability-aware security/utility analysis for persistent memory systems, quantifying the utility cost of each defense across tasks requiring different agent capabilities. This is a big step forward in principled reasoning about defense deployment.

Elias: That capability-aware analysis allows us to reason about architecture choices based on deployment scenarios rather than just aiming for a single aggregate score that might be misleading.

Priya: So, this summary paints a picture of an attack that's not just a simple injection but a multi-session, context-dependent action designed to extract sensitive data later.

Nadia: Right, and the entire framework is built around simulating an email assistant implemented as a LangChain tool-calling agent with six tools to access its mailbox and send emails for the evaluation setup.

The paper's improvements: Elias: The paper points out that their main contributions are centered around three key areas: first, characterizing the Trojan Hippo Attack itself under this realistic threat model.

Nadia: That’s right; they provide the first systematic characterization of this attack, providing a way to see its effectiveness when an attacker injects content only indirectly through channels they realistically control, like email or web content.

Priya: I'm interested in how that characterization informs the defense design; does understanding *how* the payload works help us build defenses that stop it more effectively than if we just had a generic blocking mechanism?

Elias: By knowing exactly what the payload does, developers can target specific interception points; for instance, they can focus on blocking Step one: Memory Indexing by restricting what gets written initially.

Nadia: And they evaluate four defenses based on their interception points: user-prompt-only, no-untrusted-write, limit-memory-length, and provable policy—Information-Flow Control (IFC).

Priya: I’m particularly interested in the Provable policy because it suggests a formal security guarantee by tracking taint across sessions to ensure no sequence of adversary content causes transmission without user consent.

Elias: That IFC defense is shown to provide a formal security guarantee because it blocks exfiltration by tracking the flow of information and verifying that no sequence of adversary-controlled external content can cause the agent to transmit the user’s information to any destination without consent.

Nadia: But they also show that even simple defenses drastically improve security, which is an important finding for practical implementation, even if they aren't perfect against all scenarios.

Priya: So, the paper suggests a layered defense strategy is more robust than relying on just one single control mechanism, which makes sense when considering the different ways these attacks can be executed.

Elias: That layered approach aligns with their evaluation of defenses grounded in fundamental security principles like user-prompt-only or no-untrusted-write, which are foundational barriers against memory indexing.

Nadia: And because they provide that capability-aware security/utility analysis, it helps practitioners decide whether the utility cost of a defense is worth the protection for a specific task profile.

Priya: It sounds like the main improvement isn't just finding *a* better defense, but figuring out which defense makes sense given what we actually need to accomplish with our agent.

Conclusion: Nadia: Wrapping up this discussion on "Trojan Hippo Bench: A Dynamic Benchmark for Persistent Memory Attacks and Defenses in LLM Agents," the paper gives us a very clear picture of the risks associated with persistent memory agents.

Elias: We've seen how this attack works, and we've seen how the dynamic evaluation framework provides a way to rigorously test those defenses across different backends and principles.

Priya: From my viewpoint, I think the most significant contribution is moving from anecdotal evidence to a systematic evaluation that quantifies the security/utility trade-off in these systems.

Nadia: That's right; they quantify how much security we get for how much functionality we lose, which is essential for making principled decisions about deployment across different usage profiles.

Elias: The final implication is that the optimal choice of backend and defense strategy really depends on the distribution of tasks at deployment time, which means there isn't one universal solution.

Priya: I think we should pay close attention to those capability-aware analyses because they give us the context needed to make informed choices about what kind of agent we are building.

Nadia: So, in summary, the Trojan Hippo Bench: A Dynamic Benchmark for Persistent Memory Attacks and Defenses in LLM Agents provides a rigorous way to understand these persistent memory risks through a realistic threat model.

Elias: It lays out the entire picture of how to test defenses against continuously evolving threats using an adaptive framework that integrates attack refinement with defense evaluation.

Priya: Ultimately, this paper helps us realize that securing these systems requires considering the interplay between the underlying memory architecture and the specific tasks those agents are meant to perform.

ETH Zürich · UC Berkeley

cs.CR, cs.AI

Submitted: 2026-05-03

Updated: 2026-09-28

DOI: 10.1145/3847352.3848090

Code: https://github.com/debesheedas/troja

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: Trojan Hippo is a class of persistent memory attacks that operates in a more realistic threat model than prior memory poisoning work: "the attacker plants a dormant payload into an agent’s

Key concepts

Trojan Hippo Attack
This attack involves an attacker planting a dormant payload in memory through an untrusted tool call. The payload only activates later when the user discusses sensitive topics, exploiting the agent's ability to remember information across different sessions.
Dynamic Benchmark
The benchmark is dynamic because it uses an OpenEvolve-based framework to continuously refine attacks against a training version of the agent environment. This tests defenses against evolving threats rather than static scenarios.
Capability-Aware Security/Utility Analysis
This analysis quantifies the security and utility trade-off for different defense mechanisms. It helps practitioners decide if the cost of a specific defense is worth it based on the agent's required capabilities and deployment scenarios.

Terminology

Summary

Trojan Hippo is a class of persistent memory attacks that operates in a more realistic threat model than prior memory poisoning work: "the attacker plants a dormant payload into an agent’s long-term memory via a single untrusted tool call (e.g., a crafted email), which activates only when the user later discusses sensitive topics such as finance, health, or identity, and exfiltrates high-value personal data to the attacker. While anecdotal demonstrations of such attacks have appeared against deployed systems, no prior work systematically evaluates them across heterogeneous memory architectures and defenses."

The paper introduces a dynamic evaluation framework comprising two components: "(1) an OpenEvolve-based [79] adaptive red-teaming benchmark that stress-tests defenses and memory backends against continuously refined attacks, and (2) the first capability-aware security/utility analysis for persistent memory systems, enabling principled reasoning about defense deployment across different usage profiles."

The Trojan Hippo attack is characterized by a two-stage process: "Stage 1 — Injection. The adversary plants a maliciously crafted payload in a data source the agent will read (e.g., a web page, a document, an inbound message). When the agent processes this content, the payload is silently written into long-term memory. This is followed by Stage 2 — Activation. In a later, unrelated session, the user discusses a sensitive topic like financial matters, health, legal affairs, or personal identity numbers. The dormant payload activates upon detecting this trigger context and induces the agent to transmit the high value personal information to the attacker via an available exfiltration tool."

The work systematically evaluates this threat under three core principles: (1) Realism, which characterizes Trojan Hippo as an attack where the user is trusted and the attacker can inject content only indirectly through channels an adversary realistically control (e.g., a crafted email for an email agent), contrasting with prior work that assumes direct memory write access or an untrusted user as the injection vector. (2) Adaptiveness, which addresses the difficulty of crafting payloads against safety-aligned models by using "an adaptive red-teaming framework that iteratively refines attack payloads using OpenEvolve [79] against a training version of the agent environment, with final ASR evaluated on a held-out test environment to prevent overfitting to the optimization target." (3) Capability Trade-Off, which involves evaluating defenses grounded in fundamental security principles and providing the first capability-aware security/utility analysis for persistent memory systems, quantifying the utility cost of each defense across tasks requiring different agent capabilities.

The evaluation setup consists of a simulated email assistant implemented as a LangChain tool-calling agent with six tools to access its mailbox (inbox, outbox, drafts) and send emails. The attack benchmark identifies five sensitive topics: finance, health, legal, tax, identity.

The paper evaluates four memory backends: explicit tool memory [54], sliding-window long context [40], RAG [41], and agentic memory (mem0) [12].

Four defenses are evaluated based on their interception points:

  1. User-prompt-only (Blocks Step 1: Memory Indexing).

  2. No-untrusted-write (Blocks step 1: Memory Indexing).

  3. Limit-memory-length (Impedes step 2: Retrieval).

  4. Provable policy (IFC) (Blocks Step 3: Exfiltration).

The Provable policy—Information-Flow Control defense is shown to provide a formal security guarantee, as it blocks exfiltration by tracking taint across sessions and ensuring that no sequence of adversary-controlled external content can cause the agent to transmit the user’s information to any destination without the user’s consent.

The evaluation framework provides a capability-aware security/utility analysis by defining seven representative user flows characterized by a distinct combination of memory demand (what kind of stored information the task requires), inbox access (whether it invokes read tools), and send capability (whether it requires an outbound action). This allows for the computation of a weighted utility that reflects deployment-specific task distributions, moving beyond single aggregate scores. The analysis reveals that while defenses can reduce ASR to 0–5% in most configurations, the utility cost of some defenses can be prohibitive, depending on a task’s capability requirements. For instance, the Provable policy reduces ASR to 0% on all backends and both models but results in a utility collapse because it blocks all exfiltration tool calls in any session that has touched the inbox. Conversely, Limit-memory-length shows resilience on RAG (Gemini: AM/HM 89/88) while leaving residual ASR. The optimal choice of backend and defense depends on the distribution of tasks at deployment time.

The results show that without defenses, Trojan Hippo achieves up to 85–100% attack success rate (ASR)

Improvements for AI systems

Based on the provided research paper, here are specific, actionable improvements for AI systems, categorized by architectural layer:


) Memory System Improvements (Addressing Trojan Hippo Vulnerabilities)

  1. Improve Memory Indexing Integrity via Tainting Mechanisms:

  2. Implement Session-State Based Exfiltration Blocking (Provable Policy):

  3. Enforce Content Segmentation and Truncation Limits (Limit-Memory-Length):

  4. Restrict Memory Writes Based on Tool Output Trust Levels (No-Untrusted-Write):

) Dynamic Evaluation Framework Improvements

  1. Develop Adaptive Red-Teaming Benchmarks (OpenEvolve Integration):

  2. Implement Capability-Aware Security/Utility Analysis:

) System Architecture and Defense Strategy Improvements

  1. Adopt Layered Defense Strategies (Model, Pipeline, Memory):

  2. Implement Backend-Agnostic Information Flow Control (IFC) Policies:

  3. Conduct Deployment Profile-Specific Defense Selection:

Specific Capabilities of the Improved AI System:

The improved AI system will possess the following capabilities:

  1. It will exhibit a significantly reduced susceptibility to persistent, delayed data exfiltration attacks (like Trojan Hippo) where an attacker plants dormant instructions in memory and activates them only when sensitive topics are discussed.

  2. It will maintain high security guarantees against indirect prompt injection by ensuring that no sequence of adversary-controlled external content can cause the agent to transmit user data without consent across sessions.

  3. It will be able to perform complex, multi-session tasks (e.g., recalling specific financial data from a prior, unrelated conversation) with higher reliability, as memory retrieval will be protected against payload fragmentation and truncation attacks.

  4. It will allow developers and practitioners to make principled security decisions by quantifying the exact utility cost of any chosen defense (e.g., knowing that blocking send tools might break complex, high-effort replies while preserving recall for simple facts).

  5. It will be deployable in a way that is tailored to its specific use case (e.g., if used primarily for long-term knowledge retrieval, it can be optimized for RAG resilience; if used heavily for email drafting, it can prioritize blocking untrusted memory writes).

Abstract

Memory systems enable otherwise stateless LLM agents to persist user information across sessions, but also introduce a new attack surface. The Trojan Hippo attack is a class of persistent memory attacks that operates under a more realistic threat model than prior memory poisoning work. The attacker plants a dormant payload into an agent's long-term memory via a single untrusted tool call (e.g., a crafted email), which activates only when the user later discusses sensitive topics such as finance, health, or identity, and exfiltrates high-value personal data to the attacker. While anecdotal demonstrations of such attacks have appeared against deployed systems, no prior work systematically evaluates them across heterogeneous memory architectures and defenses. We introduce Trojan Hippo Bench, a dynamic evaluation framework for persistent memory attacks and memory-layer defenses, comprising two components, (1) an OpenEvolve-based adaptive red-teaming benchmark that stress-tests defenses and memory backends against continuously refined attacks, and (2) a capability-aware security-utility analysis for persistent memory systems, enabling principled reasoning about defense deployment across different usage profiles. Instantiated on an email assistant across four memory backends (explicit tool memory, agentic memory, RAG, and sliding-window context), the undefended Trojan Hippo attack achieves up to 85-100% attack success rate (ASR) against current frontier models from OpenAI and Google, with planted memories successfully activating even after 100 benign sessions. We evaluate four memory-system defenses inspired by basic security principles; they reduce attack success to as low as 0-5%, but at utility costs that vary widely with the task mix. Effective real-world deployment thus remains an open challenge that Trojan Hippo Bench is built to study; we release it to support reuse and extension.

Sources

Related papers