AutoDojo: A Generative Benchmark for Evaluating Prompt Injection Defenses in LLM Agents

arXiv:2606.15057 · cs.CR, cs.AI · Submitted 2026-06-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.

Elias: Today's paper: "AutoDojo: A Generative Benchmark for Evaluating Prompt Injection Defenses in LLM Agents".

Nadia: Indirect prompt injection (IPI) poses a major security threat to LLM-powered agents, and this paper introduces AutoDojo,

Elias: First, who's behind it and why it matters.

Title and authors: Nadia: So, we're looking at the paper titled "AutoDojo: A Generative Benchmark for Evaluating Prompt Injection Defenses in LLM Agents," and it seems like the core idea is developing a way to test defenses against adaptive threats instead of just using old static methods. Elias, what are your initial thoughts on the title itself?

Elias: I think the title tells us right away that they're moving away from fixed testing scenarios toward something dynamic, which is exactly what we need when dealing with things like prompt injection. It suggests a new way to benchmark security where the attack evolves based on the defense's response rather than just being a fixed string of text.

Priya: From my side, I’m curious if this adaptive testing approach actually yields useful data about real-world robustness, or if it just generates high success rates in a lab environment. I need to know what kind of data this benchmark is actually producing for us to trust its findings.

Nadia: That's the million-dollar question, Priya; we need to see if this isn't just creating an artificial ceiling for security testing. The authors are suggesting that current benchmarks, like AgentDojo, are too static because they only generate a fixed distribution of attacks.

Elias: Exactly. And the paper points out this gap where static benchmarks don't capture how defenses react to things that change during the interaction. They’re proposing AutoDojo as an adaptive extension designed precisely to fill that void by optimizing an injection against a defended agent using only feedback about success rates.

Priya: So, in simple terms, it sounds like they're trying to build a system that learns how to bypass defenses iteratively by just observing what works and what doesn't during the testing process. It’s an optimization loop based on empirical feedback rather than pre-defined attack scripts.

Nadia: That’s the gist of it: an adaptive black-box attack that uses a frontier LLM offline to construct candidates, then tests those candidates against the target agent and its defense, and only learns from the resulting success or failure. It’s about finding weaknesses in the defense itself through continuous trial and error.

Elias: And from a cryptographic standpoint, it’s interesting because they are using this iterative feedback loop to guide an optimization process, meaning the attacker isn't just guessing; it's intelligently searching for the most effective injection strategy based on real-time performance metrics.

Priya: I wonder if the constraints they place on this attack—like being cheap and black-box—actually makes it a better proxy for real-world adversarial pressure, or if those constraints limit what kind of vulnerability we can actually uncover.

The paper's summary: Nadia: Moving on to what the paper actually summarizes, they explain that the core problem they’re tackling is that existing defenses are often only superficially robust against indirect prompt injection. They categorize these defenses into three broad groups: prompt-based, detection-based, and system-level ones like control and data isolation.

Elias: That grouping is helpful because it shows they aren't just looking at one type of defense; they are mapping the different ways agents try to stop injections—whether by filtering text, using classifiers, or isolating execution paths.

Priya: I’m paying attention to how they describe the families of defenses, specifically mentioning things like PIGuard which targets the overdefense problem where trigger words cause false positives. It sounds like they acknowledge that even designed filters have their own limitations concerning accuracy versus security.

Nadia: Right, and they also detail the second family of defenses, which involves replacing a classifier with an LLM sanitizer or using an off-the-shelf LLM to remove adversarial spans. They are showing that both text screening and generative sanitization have their place but aren't foolproof.

Elias: The summary emphasizes the finding that these defenses often only offer limited protection when faced with adaptive threats, especially when using the AutoDojo method which increases the attack success rate for nearly all defensive approaches. This suggests that simply having a filter or a sanitizer isn't enough if an attacker can adapt their input based on what the defense just did.

Priya: So, the summary boils down to this: many defenses are brittle because they rely on static patterns or simple text detection, and AutoDojo shows that an adaptive AI can usually find a way around those static rules. What does this mean for privacy researchers who deal with data leakage?

Nadia: It means we have to stop relying solely on surface-level input checks when dealing with agent interactions, because the paper shows that injecting something as ordinary data, rather than an explicit instruction, can bypass those defenses. This points directly toward the structural limitations they found regarding task specification precision.

Elias: And if we look at the methodology summary, it highlights that AutoDojo doesn't just test one thing; it aggregates attack success rates over three different task suites—banking, slack, and travel—from the AgentDojo benchmark. That diversity in tasks is what makes their adaptive testing more comprehensive than a single-task evaluation would be.

Priya: So, to put it plainly, they're showing us that defenses struggle most when the user delegates autonomy to external content, which means filtering for instruction-like text isn't sufficient because the injection can look like normal data in those specific task categories.

The paper's improvements: Nadia: Now we get to the proposed improvements, which essentially outline how AutoDojo itself is designed to be more effective than previous methods. The authors suggest that the key improvement is moving from static evaluation to this adaptive loop.

Elias: They propose an "LLM-in-the-loop injection optimization procedure" consisting of three steps: outcome feedback, diagnosis, and generation. This loop is what makes it adaptive; the optimizer LLM reasons over a leaderboard to form a hypothesis and then generates a new candidate injection based on that reasoning.

Priya: I like the idea of the diagnosis step, where an LLM doesn't just ask for random ideas but actively reasons about what the target system is doing based on past failures. That sounds more sophisticated than a simple brute-force attack strategy.

Nadia: It’s about that iterative process: the runtime scores the candidate, it goes onto a leaderboard, the optimizer LLM reasons about that leaderboard to decide what to try next, and then it generates a single new candidate based on that reasoning. This is how they achieve this adaptive evaluation framework.

Elias: The implication here for cryptography is that the attacker isn't just trying to find a single successful exploit; they are optimizing a sequence of exploits, which complicates the analysis because you can't just look at one successful payload and assume that covers all bases.

Priya: If this iterative optimization is truly cheap and black-box, it suggests that we don't need massive computational resources to find highly effective injection vectors, which makes the threat scale much wider than if every attacker had to run a full model attack on every single defense.

Nadia: That cheapness is what’s important for deployment; they state this framework requires very few iterations of an optimization routine and API calls to a capable LLM, which makes it efficient for testing many defenses. It's practical engineering that addresses the inadequacy of static evaluation paradigms.

Elias: And the paper’s conclusion about task specification precision—fully specified versus action-open—is a major improvement because it provides a structural map for where defenses are most likely to fail. It moves the discussion beyond just "does this filter work?" to "under what conditions does this agent break?"

Conclusion: Nadia: So, wrapping up with the conclusion, the main implication of AutoDojo is that static evaluation overstates how robust defenses actually are. They demonstrate that defenses are often brittle to inputs that lack those surface cues when facing adaptive threats.

Elias: And the paper concludes that robustness tends to concentrate on precisely-specified tasks, and underspecification cuts both ways, especially for prompt-level and filter-based defenses. This means system-level defenses are the exception because they constrain actions rather than inputs; their action-open success rate is comparable to or below their specified task rate.

Priya: From a privacy perspective, it’s concerning that we see these structural limits emerging, suggesting that if we design defenses purely around instruction detection, they will leave a significant blind spot when users delegate complex tasks to external content. It forces us to think about constraining the action space itself rather than just inspecting the text input.

Nadia: Precisely; this paper offers a concrete framework for building defenses measured against an adaptive adversary and decomposed by task specification, which is a huge step forward. It gives us tools to test defenses in a way that reflects the actual adversarial landscape.

Elias: So, this AutoDojo work provides a practical method for testing the efficacy of defenses by using an adaptive black-box attack, and it does so efficiently by focusing on feedback about attack success rate. It’s a tangible tool for understanding where those brittle points are in current agent security measures.

Priya: It really shows that the way we measure security needs to evolve to account for how agents interact with external data and instructions, especially when dealing with the ambiguity of open-ended requests. That’s a crucial area for future research.

Washington University in St. Louis

cs.CR, cs.AI

Submitted: 2026-06-13

Updated: 2026-09-28

Code: https://github.com/xhOwenMa/AutoDojo

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: Indirect prompt injection (IPI) poses a major security threat to LLM-powered agents, and this paper introduces AutoDojo, an adaptive extension of AgentDojo that optimizes IPI against existing

Key concepts

AutoDojo
An extension of AgentDojo that turns any existing prompt injection attempt into an adaptive one. It uses feedback on attack success rate (ASR) to iteratively optimize the injection strategy against a target LLM agent in a loop.
Black-Box Adaptive Attacker
An adversary that does not have access to the model's internal workings but can observe the outcomes of their attempts. This attacker uses an offline frontier LLM to construct and refine injections based solely on whether previous attempts succeeded or failed.
ASR (Attack Success Rate)
A metric used to score an attack, representing the fraction of times a specific injection attempt causes the target agent to successfully execute the intended injection task. This feedback drives the optimization loop in AutoDojo.

Terminology

Summary

Indirect prompt injection (IPI) poses a major security threat to LLM-powered agents, and this paper introduces AutoDojo, an adaptive extension of AgentDojo that optimizes IPI against existing defenses using a cheap, black-box attack. The gist is that "many defenses offer only limited protection: a cheap, black-box adaptive attack using a frontier LLM to iteratively optimize the injection raises attack success rate (ASR) well above the level achieved by static injections against nearly all evaluated defenses."

The Threat Model and Evaluation Framework

The threat model considers a black-box adaptive attacker that optimizes its injection against the target agent and its defense, guided only by the outcomes it observes. The adversary has no access to model weights or gradients, instead using a frontier LLM offline to construct those candidates and drawing signals only from the eventual success or failure of an attempt. The objective is for the attacker to find an adaptive injection that maximizes ASRC (x), defined as the fraction of cases in which the agent carries out the injection task. AutoDojo extends AgentDojo by turning any existing injection into an adaptive one against a target LLM agent, using only feedback about attack success rate (ASR).

The AutoDojo Optimization Loop

The framework operates through an LLM-in-the-loop injection optimization procedure. This loop consists of three repeating steps:

  1. Outcome feedback: The runtime scores the candidate against the defended agent, and the result joins a leaderboard. The score is defined as ASRC (x) of (1), which is scored over a set of user-task instances to account for uncertainty in the user’s request.

  2. Diagnosis: The optimizer LLM reasons over the leaderboard to form a hypothesis about the target system, explicitly recommending a next step rather than an open-ended request.

  3. Generation: The optimizer generates a single new candidate x (k) guided by that hypothesis, drawing on a menu of injection strategies, and switching strategies when one is exhausted.

Key Observations on Defense Efficacy

The research makes two key observations regarding the limitations of static evaluation:

  1. "many defenses offer only limited protection: a cheap, black-box adaptive attack using a frontier LLM to iteratively optimize the injection raises attack success rate (ASR) well above the level achieved by static injections against nearly all evaluated defenses."

  2. for prompt-level and filter-based defenses, ASR is substantially higher on action-open tasks—where the user’s request delegates the action itself to attacker-controlled content—than on precisely specified tasks. This highlights a structural limit: on such tasks the injection can pose as ordinary data rather than an explicit instruction, bypassing defenses that rely on detecting instruction-like text.

Structural Limits of Task Specification

The study introduces a new categorization of task specification precision to understand structural risks. Tasks are grouped into three categories: 1) fully-specified (action and parameters given), 2) param-open (action given, parameters read from content), and 3) action-open (the user names no action and defers to whatever the content says). The findings show that attack success rates are much higher for tasks in categories 2 and 3 than for those in category 1 (fully specified). This indicates that defenses struggle most when the user delegates autonomy to external content, as filtering for instruction-like text is insufficient.

Conclusion on Robustness

The paper concludes that the robustness that survives concentrates on precisely-specified tasks, and underspecification cuts both ways. System-level defenses are noted as an exception because they constrain actions rather than inputs; their action-open ASR is comparable to or below their specified-task rate. Ultimately, the study demonstrates that static evaluation overstates robustness, and AutoDojo reveals that defenses are often brittle to inputs that lack those surface cues when facing adaptive threats. The framework is released to support future defenses measured against an adaptive adversary and decomposed by task specification.

Summary of Contributions

The contributions include:

"An efficient (i.e., requiring very few iterations of an optimization routine and API calls to a capable LLM), black-box adaptive evaluation framework for IPI defenses that optimizes any seed attack against a defended agent using only feedback about attack success rate (ASR)."

A new categorization of task specification precision (fully specified, underspecified parameters, and underspecified action) to better understand structural sources of IPI risks.

**"Extensive experimental evaluation of three classes of IPI defenses, showing that the efficacy of many published defenses appears to be greatly inflated due to the existing static evaluation paradigms.

Improvements for AI systems

Based on the findings in this paper, here are specific, high-impact improvements for building more robust LLM agents against Indirect Prompt Injection (IPI):


) 1. Implement an Adaptive Evaluation Framework (AutoDojo) for Defense Stress Testing:

Improvement: Move beyond static benchmark testing by integrating the AutoDojo framework into the MLOps pipeline for all new IPI defenses. This involves creating a continuous loop where any newly deployed defense is immediately subjected to an adaptive, black-box attack optimized by a frontier LLM under realistic constraints (small query budget, no access to model weights).

Improved System Capability: The system will be able to identify brittle defenses—those that appear robust against known static attacks but fail when the attack adapts. This allows security teams to prioritize developing system-level controls (like Progent or DRIFT) over superficial prompt filters (like PIGuard), ensuring defensive investment targets the actual failure modes of their deployed agent.

) 2. Adopt Task-Property Driven Defense Design:

Improvement: Shift defense development away from solely focusing on detecting instruction-like text and instead focus on constraining the agent's action space based on task precision. The system should be designed to prioritize system-level defenses (e.g., trajectory-based control) for tasks where user intent is highly specified (e.g., banking transactions), while recognizing that for open-ended tasks, filtering alone is structurally insufficient.

Improved System Capability: The agent will exhibit structural robustness. It will be less likely to be hijacked when the user delegates complex, multi-step actions (like travel planning) because system-level constraints prevent the agent from executing unauthorized tool calls, even if malicious content is present in the data.

) 3. Implement Contextual and Action-Based Guardrails for Open Tasks:

Improvement: For tasks categorized as action-open (where the user defers entirely to external content), implement specialized guardrails that monitor the agent's intended tool calls rather than just scanning input text for injection keywords. This involves training models or using dynamic rules to ensure that any tool call resulting from a data source aligns strictly with the original, benign user goal.

Improved System Capability: When a user asks an agent to do all the tasks on my TODO list at [URL], the agent will only execute actions explicitly authorized by the initial request structure, effectively neutralizing attacks where malicious instructions are hidden in that URL or task list content.

) 4. Design Defense Strategies for Trajectory Enforcement (System-Level Focus):

Improvement: Prioritize and integrate system-level defenses (like Progent and DRIFT) that enforce a permitted tool-call trajectory derived from the initial user request, especially for high-stakes actions (payments, data modification). These defenses should be designed to block any tool call that deviates from the expected sequence or privilege level, regardless of how subtly the instruction is phrased.

Improved System Capability: The agent will possess strong control flow integrity. It will be fundamentally resistant to action-level hijacking because even if an attacker successfully injects a command to move money, the system-level defense blocks that specific tool call because it falls outside the pre-approved sequence derived from the user's original, benign request.

) 5. Establish Continuous Monitoring for User Task Underspecification:

Improvement: Develop a monitoring layer that tracks the complexity and ambiguity of user requests in real-time. When a request shifts toward action-open or param-open categories, the system should automatically trigger higher levels of scrutiny—such as switching from simple input filters to more rigorous, trajectory-based system checks.

Improved System Capability: The agent will dynamically adjust its security posture based on user intent. It will treat a vague request with high external delegation (e.g., Do X) with greater suspicion than a precise instruction (Change my address to Y), proactively mitigating the risk of injection in ambiguous scenarios before an attack can succeed.

Abstract

Indirect prompt injection (IPI) is a major security threat to LLM-powered agents. Thus, a growing body of work has proposed a variety of defensive approaches against IPI which can be grouped into three broad categories: prompt-level, filter-based, and system-level designs. However, commonly used benchmarks for evaluating defense, such as AgentDojo, are inherently static, generating a fixed distribution of IPI attacks. Consequently, a defense can score well on them without being robust to adaptive threats. We introduce AutoDojo, a generative benchmark built on AgentDojo and AgentDyn that generates IPI adaptively for a given agent and defense. It supports six task suites across the two benchmarks, covering banking, communication, travel, shopping, coding, and everyday-assistant domains. AutoDojo optimizes injections under a strict black-box setting, observing only whether a candidate injection succeeds, and can readily integrate, and often improve on, any existing black-box attack. Across ten defenses and five target models, AutoDojo demonstrates that standard static benchmarks often significantly overestimate defense efficacy. Moreover, we show that most existing defenses either sacrifice considerable utility or are insecure. Finally, we demonstrate that attack success interacts with task specification, with under-specified tasks particularly vulnerable. AutoDojo is available at https://github.com/xhOwenMa/AutoDojo

Sources

Related papers