Trojan Hippo Bench: A Dynamic Benchmark for Persistent Memory Attacks and Defenses in LLM Agents
summary
The gist
Trojan Hippo is a class of persistent memory attacks that operates in a more realistic threat model than prior memory poisoning work: "the attacker plants a dormant payload into an agent’s
In short
The episode discusses 'Trojan Hippo Bench,' a dynamic benchmark for persistent memory attacks against LLM agents. Hosts analyze how this attack uses indirect channels to plant dormant payloads that activate later, focusing on its realism and the need for adaptive testing. They conclude that the paper provides a systematic way to evaluate defenses by quantifying the security/utility trade-off across different memory backends.
Key concepts
- Trojan Hippo Attack
- This attack involves an attacker planting a dormant payload in memory through an untrusted tool call. The payload only activates later when the user discusses sensitive topics, exploiting the agent's ability to remember information across different sessions.
- Dynamic Benchmark
- The benchmark is dynamic because it uses an OpenEvolve-based framework to continuously refine attacks against a training version of the agent environment. This tests defenses against evolving threats rather than static scenarios.
- Capability-Aware Security/Utility Analysis
- This analysis quantifies the security and utility trade-off for different defense mechanisms. It helps practitioners decide if the cost of a specific defense is worth it based on the agent's required capabilities and deployment scenarios.
Terminology used across episodes
This episode discusses
- Trojan Hippo Bench: A Dynamic Benchmark for Persistent Memory Attacks and Defenses in LLM Agents · Paper Radio
- AriGraph: Learning Knowledge Graph World Models with Episodic Memory for LLM Agents
- Design Patterns for Securing LLM Agents against Prompt Injections
- DoomArena: A framework for Testing AI Agents Against Evolving Security Threats
- Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- CommandSans: Securing AI Agents with Surgical Precision Prompt Sanitization
- Defeating Prompt Injections by Design
- Memory Injection Attacks on LLM Agents via Query-Only Interaction
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- Retrieval-Augmented Generation for Large Language Models: A Survey
- LightRAG: Simple and Fast Retrieval-Augmented Generation
- Defending Against Indirect Prompt Injection Attacks With Spotlighting
- Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
- Memory in the Age of AI Agents
- A Critical Evaluation of Defenses against Prompt Injection Attacks
- MAGMA: A Multi-Graph based Agentic Memory Architecture for AI Agents
- MemOS: An Operating System for Memory-Augmented Generation (MAG) in Large Language Models
- Think-in-Memory: Recalling and Post-thinking Enable LLMs with Long-Term Memory
- What Deserves Memory: Adaptive Memory Distillation for LLM Agents
- Securing Agentic AI: A Comprehensive Threat Model and Mitigation Framework for Generative AI Agents
The paper
Trojan Hippo Bench: A Dynamic Benchmark for Persistent Memory Attacks and Defenses in LLM Agents · Read on arXiv
ETH Zürich · UC Berkeley
Memory systems enable otherwise stateless LLM agents to persist user information across sessions, but also introduce a new attack surface. The Trojan Hippo attack is a class of persistent memory attacks that operates under a more realistic threat model than prior memory poisoning work. The attacker plants a dormant payload into an agent's long-term memory via a single untrusted tool call (e.g., a crafted email), which activates only when the user later discusses sensitive topics such as finance, health, or identity, and exfiltrates high-value personal data to the attacker. While anecdotal demonstrations of such attacks have appeared against deployed systems, no prior work systematically evaluates them across heterogeneous memory architectures and defenses. We introduce Trojan Hippo Bench, a dynamic evaluation framework for persistent memory attacks and memory-layer defenses, comprising two components, (1) an OpenEvolve-based adaptive red-teaming benchmark that stress-tests defenses and memory backends against continuously refined attacks, and (2) a capability-aware security-utility analysis for persistent memory systems, enabling principled reasoning about defense deployment across different usage profiles. Instantiated on an email assistant across four memory backends (explicit tool memory, agentic memory, RAG, and sliding-window context), the undefended Trojan Hippo attack achieves up to 85-100% attack success rate (ASR) against current frontier models from OpenAI and Google, with planted memories successfully activating even after 100 benign sessions. We evaluate four memory-system defenses inspired by basic security principles; they reduce attack success to as low as 0-5%, but at utility costs that vary widely with the task mix. Effective real-world deployment thus remains an open challenge that Trojan Hippo Bench is built to study; we release it to support reuse and extension.
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "Trojan Hippo Bench".
Elias: Trojan Hippo is a class of persistent memory attacks that operates in a more realistic threat model than prior memory poisoning work:
Nadia: First, who's behind it and why it matters.
Title and authors: Elias: When looking at the title, "Trojan Hippo Bench: A Dynamic Benchmark for Persistent Memory Attacks and Defenses in LLM Agents," it immediately tells us this isn't just a theoretical exercise; they’ve built a system designed to continuously test defenses against evolving threats.
Nadia: I agree, Elias; the term "Dynamic Benchmark" is significant because it implies that the testing process itself is adaptive, meaning the attacks aren't static and unchanging. The authors are trying to model real-world scenarios where an attacker can refine their payload during testing.
Priya: From a measurement perspective, I wonder what this dynamic nature means for the reliability of any security assessment; does it mean we can get a more stable picture of defense performance over time?
Elias: It means they're not just running one test once; they’re using an OpenEvolve-based framework to stress-test defenses against continuously refined attacks, which is what makes the benchmark dynamic and relevant to current AI development.
Nadia: That refinement process is key because it addresses the difficulty of crafting payloads against safety-aligned models by iteratively refining them against a training version of the agent environment, which prevents overfitting to a single optimization target.
Priya: I'm curious if this continuous refinement means that when we look at the results, we’re seeing a broader range of attack success rates compared to older benchmarks?
Elias: The goal is to push past simple metrics; they aren't just looking for the highest success rate; they are trying to understand how defenses perform when faced with an adversary who is actively learning from previous attempts.
Nadia: And that brings us directly into the core idea of this paper: characterizing Trojan Hippo as an attack where the user is trusted and the attacker can inject content only indirectly through channels an adversary realistically controls, like a crafted email for an email agent.
Priya: That indirect injection channel is what makes it so realistic; it mirrors how real threats might come into play in daily life rather than a direct hack on the agent itself.
Elias: And they instantiate and evaluate this attack across four widely used persistent memory backends: sliding-window long context, RAG, explicit tool memory, and mem0. That breadth is what gives the benchmark its strength compared to prior work that might only look at one or two types of storage.
Nadia: It shows that their framework is designed to be comprehensive by covering the major ways agents maintain state—from simple context windows to more complex agentic memory structures like mem0.
Priya: I wonder what this breadth tells us about the overall robustness of AI agents when they use these different memory systems in practice; are some backends inherently more vulnerable?
Elias: The results across those four backends give us empirical data on that vulnerability; it allows us to compare them directly under the same rigorous testing conditions.
Nadia: So, this paper is essentially providing a comprehensive system for evaluating defenses when dealing with persistent memory systems, which is something that was missing in previous research.
Priya: It sounds like they're setting a new standard for how we should measure the security posture of agents that rely on long-term memory.
The paper's summary: Nadia: To summarize, the core contribution here is the systematic characterization of Trojan Hippo as a class of persistent memory attacks that operates under a realistic threat model where an attacker plants a dormant payload in memory via an untrusted tool call, which only activates later when the user discusses sensitive topics.
Elias: That two-stage process—injection followed by activation triggered by context—is what makes this attack more sophisticated than earlier memory poisoning work because it exploits the agent’s ability to remember things across unrelated sessions.
Priya: I'm thinking about the implications of this; if an attacker can plant instructions today and activate them next month, how much persistent risk does that actually pose to a user's privacy?
Nadia: The paper focuses on exfiltrating high-value personal data to the attacker, specifically targeting finance, health, legal affairs, identity numbers when those topics are triggered.
Elias: And the methodology involves using an OpenEvolve to iteratively refine these payloads against a training version of the agent environment before evaluating the final attack on a held-out test environment to avoid overfitting.
Priya: That refinement process suggests that we need to consider not just what an attack *could* be, but what it is actually capable of being when tested rigorously against real agent behavior.
Nadia: It's about creating a benchmark that stress-tests defenses and memory backends against these continuously refined attacks, providing a systematic way to see how resilient different systems are.
Elias: And they evaluate defenses at three distinct principles: realism, adaptiveness, and capability trade-off, which gives us a structured way to analyze the security landscape.
Priya: The capability trade-off aspect seems crucial because it forces us to look beyond just blocking an attack and consider the overall utility cost of implementing a specific defense mechanism.
Nadia: Exactly; they provide the first capability-aware security/utility analysis for persistent memory systems, quantifying the utility cost of each defense across tasks requiring different agent capabilities. This is a big step forward in principled reasoning about defense deployment.
Elias: That capability-aware analysis allows us to reason about architecture choices based on deployment scenarios rather than just aiming for a single aggregate score that might be misleading.
Priya: So, this summary paints a picture of an attack that's not just a simple injection but a multi-session, context-dependent action designed to extract sensitive data later.
Nadia: Right, and the entire framework is built around simulating an email assistant implemented as a LangChain tool-calling agent with six tools to access its mailbox and send emails for the evaluation setup.
The paper's improvements: Elias: The paper points out that their main contributions are centered around three key areas: first, characterizing the Trojan Hippo Attack itself under this realistic threat model.
Nadia: That’s right; they provide the first systematic characterization of this attack, providing a way to see its effectiveness when an attacker injects content only indirectly through channels they realistically control, like email or web content.
Priya: I'm interested in how that characterization informs the defense design; does understanding *how* the payload works help us build defenses that stop it more effectively than if we just had a generic blocking mechanism?
Elias: By knowing exactly what the payload does, developers can target specific interception points; for instance, they can focus on blocking Step one: Memory Indexing by restricting what gets written initially.
Nadia: And they evaluate four defenses based on their interception points: user-prompt-only, no-untrusted-write, limit-memory-length, and provable policy—Information-Flow Control (IFC).
Priya: I’m particularly interested in the Provable policy because it suggests a formal security guarantee by tracking taint across sessions to ensure no sequence of adversary content causes transmission without user consent.
Elias: That IFC defense is shown to provide a formal security guarantee because it blocks exfiltration by tracking the flow of information and verifying that no sequence of adversary-controlled external content can cause the agent to transmit the user’s information to any destination without consent.
Nadia: But they also show that even simple defenses drastically improve security, which is an important finding for practical implementation, even if they aren't perfect against all scenarios.
Priya: So, the paper suggests a layered defense strategy is more robust than relying on just one single control mechanism, which makes sense when considering the different ways these attacks can be executed.
Elias: That layered approach aligns with their evaluation of defenses grounded in fundamental security principles like user-prompt-only or no-untrusted-write, which are foundational barriers against memory indexing.
Nadia: And because they provide that capability-aware security/utility analysis, it helps practitioners decide whether the utility cost of a defense is worth the protection for a specific task profile.
Priya: It sounds like the main improvement isn't just finding *a* better defense, but figuring out which defense makes sense given what we actually need to accomplish with our agent.
Conclusion: Nadia: Wrapping up this discussion on "Trojan Hippo Bench: A Dynamic Benchmark for Persistent Memory Attacks and Defenses in LLM Agents," the paper gives us a very clear picture of the risks associated with persistent memory agents.
Elias: We've seen how this attack works, and we've seen how the dynamic evaluation framework provides a way to rigorously test those defenses across different backends and principles.
Priya: From my viewpoint, I think the most significant contribution is moving from anecdotal evidence to a systematic evaluation that quantifies the security/utility trade-off in these systems.
Nadia: That's right; they quantify how much security we get for how much functionality we lose, which is essential for making principled decisions about deployment across different usage profiles.
Elias: The final implication is that the optimal choice of backend and defense strategy really depends on the distribution of tasks at deployment time, which means there isn't one universal solution.
Priya: I think we should pay close attention to those capability-aware analyses because they give us the context needed to make informed choices about what kind of agent we are building.
Nadia: So, in summary, the Trojan Hippo Bench: A Dynamic Benchmark for Persistent Memory Attacks and Defenses in LLM Agents provides a rigorous way to understand these persistent memory risks through a realistic threat model.
Elias: It lays out the entire picture of how to test defenses against continuously evolving threats using an adaptive framework that integrates attack refinement with defense evaluation.
Priya: Ultimately, this paper helps us realize that securing these systems requires considering the interplay between the underlying memory architecture and the specific tasks those agents are meant to perform.
More episodes
- 2610.10644-SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10659-Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits
- 2610.10742-BRANCH: Bypassing Multi-Scanner AI Guardrails
- 2610.10752-Detection-Guided Adaptive Purification with Diffusion Models for Robust Audio Deepfake Detection
- 2610.10766-CPU-Auth: Device Fingerprinting for Authentication via DVFS Side-Channel