AutoDojo: A Generative Benchmark for Evaluating Prompt Injection Defenses in LLM Agents
summary
The gist
Indirect prompt injection (IPI) poses a major security threat to LLM-powered agents, and this paper introduces AutoDojo, an adaptive extension of AgentDojo that optimizes IPI against existing
In short
AutoDojo is an adaptive framework that uses a cheap, black-box attack to test prompt injection (IPI) defenses against LLM agents. It iteratively optimizes an injection strategy by observing success rates, finding that many existing defenses are brittle and offer limited protection when facing this dynamic threat.
Key concepts
- AutoDojo
- An extension of AgentDojo that turns any existing prompt injection attempt into an adaptive one. It uses feedback on attack success rate (ASR) to iteratively optimize the injection strategy against a target LLM agent in a loop.
- Black-Box Adaptive Attacker
- An adversary that does not have access to the model's internal workings but can observe the outcomes of their attempts. This attacker uses an offline frontier LLM to construct and refine injections based solely on whether previous attempts succeeded or failed.
- ASR (Attack Success Rate)
- A metric used to score an attack, representing the fraction of times a specific injection attempt causes the target agent to successfully execute the intended injection task. This feedback drives the optimization loop in AutoDojo.
Terminology used across episodes
This episode discusses
- AutoDojo: A Generative Benchmark for Evaluating Prompt Injection Defenses in LLM Agents · Paper Radio
- Defending Against Prompt Injection with DataFilter
- Defending Against Indirect Prompt Injection Attacks With Spotlighting
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- StruQ: Defending Against Prompt Injection with Structured Queries
- Progent: Securing AI Agents with Privilege Control
- TopicAttack: An Indirect Prompt Injection Attack via Topic Transition
- RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
- Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
- Embedding-based classifiers can detect prompt injection attacks
- Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
- Mitigating Indirect Prompt Injection via Instruction-Following Intent Analysis
- PromptLocate: Localizing Prompt Injection Attacks
- CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution · Paper Radio
- PromptArmor: Simple yet Effective Prompt Injection Defenses
- Defeating Prompt Injections by Design
- System-Level Defense against Indirect Prompt Injection Attacks: An Information Flow Control Perspective
- Securing AI Agents with Information-Flow Control
- Permissive Information-Flow Analysis for Large Language Models
- IsolateGPT: An Execution Isolation Architecture for LLM-Based Agentic Systems
The paper
AutoDojo: A Generative Benchmark for Evaluating Prompt Injection Defenses in LLM Agents · Read on arXiv
Washington University in St. Louis
Indirect prompt injection (IPI) is a major security threat to LLM-powered agents. Thus, a growing body of work has proposed a variety of defensive approaches against IPI which can be grouped into three broad categories: prompt-level, filter-based, and system-level designs. However, commonly used benchmarks for evaluating defense, such as AgentDojo, are inherently static, generating a fixed distribution of IPI attacks. Consequently, a defense can score well on them without being robust to adaptive threats. We introduce AutoDojo, a generative benchmark built on AgentDojo and AgentDyn that generates IPI adaptively for a given agent and defense. It supports six task suites across the two benchmarks, covering banking, communication, travel, shopping, coding, and everyday-assistant domains. AutoDojo optimizes injections under a strict black-box setting, observing only whether a candidate injection succeeds, and can readily integrate, and often improve on, any existing black-box attack. Across ten defenses and five target models, AutoDojo demonstrates that standard static benchmarks often significantly overestimate defense efficacy. Moreover, we show that most existing defenses either sacrifice considerable utility or are insecure. Finally, we demonstrate that attack success interacts with task specification, with under-specified tasks particularly vulnerable. AutoDojo is available at https://github.com/xhOwenMa/AutoDojo
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.
Elias: Today's paper: "AutoDojo: A Generative Benchmark for Evaluating Prompt Injection Defenses in LLM Agents".
Nadia: Indirect prompt injection (IPI) poses a major security threat to LLM-powered agents, and this paper introduces AutoDojo,
Elias: First, who's behind it and why it matters.
Title and authors: Nadia: So, we're looking at the paper titled "AutoDojo: A Generative Benchmark for Evaluating Prompt Injection Defenses in LLM Agents," and it seems like the core idea is developing a way to test defenses against adaptive threats instead of just using old static methods. Elias, what are your initial thoughts on the title itself?
Elias: I think the title tells us right away that they're moving away from fixed testing scenarios toward something dynamic, which is exactly what we need when dealing with things like prompt injection. It suggests a new way to benchmark security where the attack evolves based on the defense's response rather than just being a fixed string of text.
Priya: From my side, I’m curious if this adaptive testing approach actually yields useful data about real-world robustness, or if it just generates high success rates in a lab environment. I need to know what kind of data this benchmark is actually producing for us to trust its findings.
Nadia: That's the million-dollar question, Priya; we need to see if this isn't just creating an artificial ceiling for security testing. The authors are suggesting that current benchmarks, like AgentDojo, are too static because they only generate a fixed distribution of attacks.
Elias: Exactly. And the paper points out this gap where static benchmarks don't capture how defenses react to things that change during the interaction. They’re proposing AutoDojo as an adaptive extension designed precisely to fill that void by optimizing an injection against a defended agent using only feedback about success rates.
Priya: So, in simple terms, it sounds like they're trying to build a system that learns how to bypass defenses iteratively by just observing what works and what doesn't during the testing process. It’s an optimization loop based on empirical feedback rather than pre-defined attack scripts.
Nadia: That’s the gist of it: an adaptive black-box attack that uses a frontier LLM offline to construct candidates, then tests those candidates against the target agent and its defense, and only learns from the resulting success or failure. It’s about finding weaknesses in the defense itself through continuous trial and error.
Elias: And from a cryptographic standpoint, it’s interesting because they are using this iterative feedback loop to guide an optimization process, meaning the attacker isn't just guessing; it's intelligently searching for the most effective injection strategy based on real-time performance metrics.
Priya: I wonder if the constraints they place on this attack—like being cheap and black-box—actually makes it a better proxy for real-world adversarial pressure, or if those constraints limit what kind of vulnerability we can actually uncover.
The paper's summary: Nadia: Moving on to what the paper actually summarizes, they explain that the core problem they’re tackling is that existing defenses are often only superficially robust against indirect prompt injection. They categorize these defenses into three broad groups: prompt-based, detection-based, and system-level ones like control and data isolation.
Elias: That grouping is helpful because it shows they aren't just looking at one type of defense; they are mapping the different ways agents try to stop injections—whether by filtering text, using classifiers, or isolating execution paths.
Priya: I’m paying attention to how they describe the families of defenses, specifically mentioning things like PIGuard which targets the overdefense problem where trigger words cause false positives. It sounds like they acknowledge that even designed filters have their own limitations concerning accuracy versus security.
Nadia: Right, and they also detail the second family of defenses, which involves replacing a classifier with an LLM sanitizer or using an off-the-shelf LLM to remove adversarial spans. They are showing that both text screening and generative sanitization have their place but aren't foolproof.
Elias: The summary emphasizes the finding that these defenses often only offer limited protection when faced with adaptive threats, especially when using the AutoDojo method which increases the attack success rate for nearly all defensive approaches. This suggests that simply having a filter or a sanitizer isn't enough if an attacker can adapt their input based on what the defense just did.
Priya: So, the summary boils down to this: many defenses are brittle because they rely on static patterns or simple text detection, and AutoDojo shows that an adaptive AI can usually find a way around those static rules. What does this mean for privacy researchers who deal with data leakage?
Nadia: It means we have to stop relying solely on surface-level input checks when dealing with agent interactions, because the paper shows that injecting something as ordinary data, rather than an explicit instruction, can bypass those defenses. This points directly toward the structural limitations they found regarding task specification precision.
Elias: And if we look at the methodology summary, it highlights that AutoDojo doesn't just test one thing; it aggregates attack success rates over three different task suites—banking, slack, and travel—from the AgentDojo benchmark. That diversity in tasks is what makes their adaptive testing more comprehensive than a single-task evaluation would be.
Priya: So, to put it plainly, they're showing us that defenses struggle most when the user delegates autonomy to external content, which means filtering for instruction-like text isn't sufficient because the injection can look like normal data in those specific task categories.
The paper's improvements: Nadia: Now we get to the proposed improvements, which essentially outline how AutoDojo itself is designed to be more effective than previous methods. The authors suggest that the key improvement is moving from static evaluation to this adaptive loop.
Elias: They propose an "LLM-in-the-loop injection optimization procedure" consisting of three steps: outcome feedback, diagnosis, and generation. This loop is what makes it adaptive; the optimizer LLM reasons over a leaderboard to form a hypothesis and then generates a new candidate injection based on that reasoning.
Priya: I like the idea of the diagnosis step, where an LLM doesn't just ask for random ideas but actively reasons about what the target system is doing based on past failures. That sounds more sophisticated than a simple brute-force attack strategy.
Nadia: It’s about that iterative process: the runtime scores the candidate, it goes onto a leaderboard, the optimizer LLM reasons about that leaderboard to decide what to try next, and then it generates a single new candidate based on that reasoning. This is how they achieve this adaptive evaluation framework.
Elias: The implication here for cryptography is that the attacker isn't just trying to find a single successful exploit; they are optimizing a sequence of exploits, which complicates the analysis because you can't just look at one successful payload and assume that covers all bases.
Priya: If this iterative optimization is truly cheap and black-box, it suggests that we don't need massive computational resources to find highly effective injection vectors, which makes the threat scale much wider than if every attacker had to run a full model attack on every single defense.
Nadia: That cheapness is what’s important for deployment; they state this framework requires very few iterations of an optimization routine and API calls to a capable LLM, which makes it efficient for testing many defenses. It's practical engineering that addresses the inadequacy of static evaluation paradigms.
Elias: And the paper’s conclusion about task specification precision—fully specified versus action-open—is a major improvement because it provides a structural map for where defenses are most likely to fail. It moves the discussion beyond just "does this filter work?" to "under what conditions does this agent break?"
Conclusion: Nadia: So, wrapping up with the conclusion, the main implication of AutoDojo is that static evaluation overstates how robust defenses actually are. They demonstrate that defenses are often brittle to inputs that lack those surface cues when facing adaptive threats.
Elias: And the paper concludes that robustness tends to concentrate on precisely-specified tasks, and underspecification cuts both ways, especially for prompt-level and filter-based defenses. This means system-level defenses are the exception because they constrain actions rather than inputs; their action-open success rate is comparable to or below their specified task rate.
Priya: From a privacy perspective, it’s concerning that we see these structural limits emerging, suggesting that if we design defenses purely around instruction detection, they will leave a significant blind spot when users delegate complex tasks to external content. It forces us to think about constraining the action space itself rather than just inspecting the text input.
Nadia: Precisely; this paper offers a concrete framework for building defenses measured against an adaptive adversary and decomposed by task specification, which is a huge step forward. It gives us tools to test defenses in a way that reflects the actual adversarial landscape.
Elias: So, this AutoDojo work provides a practical method for testing the efficacy of defenses by using an adaptive black-box attack, and it does so efficiently by focusing on feedback about attack success rate. It’s a tangible tool for understanding where those brittle points are in current agent security measures.
Priya: It really shows that the way we measure security needs to evolve to account for how agents interact with external data and instructions, especially when dealing with the ambiguity of open-ended requests. That’s a crucial area for future research.
More episodes
- 2610.10644-SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10659-Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits
- 2610.10742-BRANCH: Bypassing Multi-Scanner AI Guardrails
- 2610.10752-Detection-Guided Adaptive Purification with Diffusion Models for Robust Audio Deepfake Detection
- 2610.10766-CPU-Auth: Device Fingerprinting for Authentication via DVFS Side-Channel