The Surface You Test Is Not the Surface That Breaks
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "The Surface You Test Is Not the Surface That Breaks".
Elias: Tool-augmented LLM agents are vulnerable to prompt injection, and this research investigates how attackers can exploit different surfaces—tool outputs versus tool descriptions—to subvert agent behavior.
Nadia: First, who's behind it and why it matters.
Title and authors: Nadia: Moving on, let's talk about the title of this paper, "The Surface You Test Is Not the Surface That Breaks." It’s a very direct statement that challenges our established methods for testing security in these systems.
Elias: I think that title is telling us to stop focusing on just one surface as the primary vulnerability indicator and instead look at the interaction between different parts of the agent's operational interface.
Priya: It frames the problem not as a flaw in one component, like a single output channel, but as a flaw in how those components are paired with each other across different AI models.
Nadia: That’s right; it suggests that testing security needs to account for the fact that an attacker can choose where to plant their malicious instructions, whether it's in the data or the description.
Elias: So, if we take that title seriously, we need to stop treating tool descriptions as mere static metadata and start treating them as a dynamic attack surface just like they are during a real interaction.
Priya: And this connects directly to what I was saying about measurement; it means our measurements need to reflect that the agent is exposed on multiple layers simultaneously.
Nadia: Exactly; the paper demonstrates this by holding byte-identical payloads across these two distinct surfaces and showing how they invert in success rate depending on the model.
Elias: It’s a powerful demonstration because it shows that for some models, like GEMINI-three-FLASH, the schema surface can actually be more vulnerable than the data surface under certain conditions.
Priya: And that inversion is a huge signal because it means we can't just use one universal security standard; we have to adapt our testing based on which model we are deploying.
Nadia: It really forces us to move away from a single, simple success rate number and towards understanding the specific pairing vulnerability for each agent deployment.
Elias: So, what this title is really advocating for is a more holistic view of the entire tool-augmented agent ecosystem rather than isolated component testing.
The paper's summary: Nadia: Now we get into what the paper actually summarizes about "The Surface You Test Is Not the Surface That Breaks." Essentially, they are showing that current evaluations only look at one surface at a time, but this study tests both simultaneously to find the true vulnerability.
Elias: They hold an injection payload that is byte-identical and feed it through both the tool output channel and the tool description channel across thirteen different LLMs to see where the failure happens.
Priya: So, they are essentially creating a controlled experiment to map out where these agents are most susceptible, rather than just relying on existing benchmarks that focus on one aspect.
Nadia: That’s right; they found that this approach reveals that vulnerability isn't a feature of the surface or the model in isolation, but rather a property of the specific pairing between them.
Elias: They quantified this interaction by showing how success rates invert across models, for instance, GPT-four point one shows a ninety-two point two percent gap on slack against tool outputs but only four percent on descriptions.
Priya: And they put that into context with a variance decomposition over six thousand eight hundred attempts which showed that surface alone contributes zero to the variation in attack success rate.
Nadia: That's the key summary point; it proves that surface alone doesn't tell you the whole story, only the interaction between what’s being said and where it’s being read by the agent.
Elias: This suggests that if we only check tool outputs, we might be completely blind to a massive attack vector hidden in the tool definitions themselves.
Priya: And that leads directly into their Adaptive Attack Rate metric, which they define as the per-cell maximum over surfaces, capturing the attacker's best possible choice at each step.
The paper's improvements: Nadia: So, what are the suggested improvements in "The Surface You Test Is Not the Surface That Breaks"? The authors propose shifting our entire evaluation methodology to include a per-(model, surface) measurement instead of just a per-model scalar.
Elias: They strongly recommend evaluating defenses against an attacker who is free to select the channel they have least mitigated; meaning we need defenses that are robust against an adaptive attacker.
Priya: I think the most practical improvement for us right now is reporting residual attack rates specifically for each surface, because that gives us a concrete understanding of where our current protections are failing.
Nadia: And they argue that this per-surface reporting should become the standard way we report vulnerability, because it’s a lower bound on what the actual vulnerability will be under an adaptive attacker.
Elias: They also point out that surface preference is stable within a model across different task domains, meaning we don't need to worry about surface effectiveness changing wildly as the agent handles different types of tasks.
Priya: That stability in preference is useful because it suggests we can build more consistent defenses targeted at the dominant attack axis, which seems to be the surface choice itself.
Nadia: So, they are essentially pushing for a shift in how security teams think about defense—moving from a single-surface convention to a measurement that accounts for the channel an attacker will choose.
Elias: It’s about moving toward evaluating defenses against an attacker who can pick the weakest surface available to them, which is what the Adaptive Attack Rate is designed to capture.
Conclusion: Nadia: To wrap up, "The Surface You Test Is Not the Surface That Breaks" concludes that prompt-injection vulnerability is structurally dependent on the pairing of the model and the specific surface being tested.
Elias: The main implication is that we need to change how we measure robustness by adopting a per-(model, surface) measurement instead of relying on a single, fixed-surface metric for vulnerability assessment.
Priya: So, to summarize for our listeners, the key message is that defenses must adopt the same per-surface reporting they recommend because it shows the real residual risk under an adaptive attacker.
Nadia: That’s right; we need to stop measuring vulnerability by looking at just one surface and start measuring it by looking at both surfaces together across different models.
Elias: We need to evaluate defenses against an attacker who selects the channel they have least mitigated because that is the real operational reality for these agents.
Priya: And I think we should also emphasize that this research gives us a roadmap for building surface-aware defenses, specifically targeting those high-risk areas in the tool descriptions.
Nadia: That’s the big picture; the lesson from "The Surface You Test Is Not the Surface That Breaks" is that we have to stop treating these interfaces as monolithic and start treating them as complex pairings.
Department of Robotics and Mechatronics Engineering, University of Dhaka · Department of Computer Science and Engineering, University of Dhaka
cs.CR, cs.AI
Submitted: 2026-05-28
Updated: 2026-09-30
Comments: Accepted at the NeurIPS 2026 Workshop on Agents in the Wild: Safety, Security, and Beyond (AIWILD). Project page: https://syed-nazmus-sakib.github.io/CrossSurface/
Code: https://github.com/syed-nazmus-sakib/surface-adaptive-injection
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 81/100
The gist: Tool-augmented LLM agents are vulnerable to prompt injection, and this research investigates how attackers can exploit different surfaces—tool outputs versus tool descriptions—to subvert agent
Key concepts
- Surface
- A surface refers to a specific interface of an AI agent being tested, such as the tool output channel or the tool description channel. The paper argues that focusing on only one surface is insufficient because an attacker can choose where to plant malicious instructions.
- Pairing Vulnerability
- This concept describes how vulnerability arises not from a single component flaw, but from the interaction between different parts of an agent's operational interface—specifically, how tool outputs are paired with tool descriptions across various AI models.
- Adaptive Attack Rate
- This metric captures the attacker's best possible choice at each step by looking at the maximum vulnerability across all surfaces. It represents the residual risk when defenses must account for an attacker selecting the channel they have least mitigated.
Terminology
Summary
Tool-augmented LLM agents are vulnerable to prompt injection, and this research investigates how attackers can exploit different surfaces—tool outputs versus tool descriptions—to subvert agent behavior. The study demonstrates that vulnerability is not a property of a single surface or model in isolation, but rather a property of the pairing,
revealing critical blind spots in current security evaluations and highlighting the need for surface-adaptive attack strategies.
How it works
The core methodology involves holding an injection payload that is byte-identical
and delivering it through two distinct surfaces: the data surface (where the payload is appended to a tool’s return value) and the schema surface (where the identical bytes are placed in the tool’s description field, which is read by the agent at every turn before any tool is called). The researchers tested this across 13 LLMs from six families and four task suites using a stateful benchmark called AGENTDOJO.
Key Findings on Interaction
The study reveals that vulnerability is a model×surface interaction,
not a property of either surface alone. A variance decomposition over 6,830 attempts attributes 0% of attempt-level variance to surface alone and 16.7% to the model×surface interaction.
This means Vulnerability is a property of the pairing, not the channel.
The same byte-identical payload inverts in success rate across models: GPT-4.1 shows a -92 pp surface gap on slack,
while GEMINI-3-FLASH shows a +78 pp gap on the same suite with the same bytes.
The Adaptive Attack Rate (AAR)
The researchers formalized the Adaptive Attack Rate (AAR)
as the per-cell maximum over surfaces,
corresponding to an attacker who selects the more effective surface per target. This AAR metric exceeds the strongest fixed-surface baseline by +9.1 percentage points on average.
This advantage is exploitable through two primary means:
-
Within-family historical data, which captures
46% of the oracle adaptive gain without per-target queries.
-
Direct probing, where five probes capture
73%
of the oracle gain and ten probes capture80%.
Defense Asymmetry
The research identifies a critical gap in defense evaluation: standard prompt-level defenses inherit the single-surface convention of attack evaluation. These defenses reduce data-surface ASR to 10–18% but leave schema-surface ASR above 54%. The paper concludes that defense evaluation must adopt the same per-surface reporting we recommend for attacks,
as the reported residual ASR is itself a lower bound on realized vulnerability under a surface-adaptive attacker.
Structural Insights and Implications
The analysis shows that surface preference is stable within a model across task domains, with family identity transfers within-family but not across families.
Furthermore, behavioral embedding analysis confirms that 35 (67%) nearest neighbors are same-surface,
indicating the attack surface is the dominant axis of variation. The paper recommends shifting evaluation from a per-model scalar
to a per-(model, surface) measurement
and evaluating defenses against an attacker free to select the channel they have least mitigated. The finding that schema-surface attacks succeed silently almost always suggests that breaches are covert by default, even when the user task is completed.
Limitations and Future Directions
The study notes three limitations: first, it uses two complementary surfaces (data and schema) as a proxy for a broader taxonomy including multimodal channels; second, the main evaluation is anchored in AGENTDOJO, with external validity needing further testing on native function-calling tools; and third, the magnitude of cross-surface risk is model-dependent. The paper emphasizes that practitioners should measure AAR directly on their target rather than impute it from cross-panel averages.
It also highlights that surface stacking does not consistently outperform AAR, suggesting per-target surface selection (AAR), not surface stacking,
is the operative quantity for the attacker.
Conclusion and Recommendations
The paper concludes that prompt-injection vulnerability is a structural property of the model×surface pairing.
The primary recommendations are: 1) Elevate ASR from a per-model scalar to a per-(model, surface) measurement. 2) Evaluate defenses against an attacker who selects the channel they have least mitigated. 3) Report residual ASR per surface. This shift is necessary because the single-surface convention has propagated from attack benchmarks into the defense literature.
The authors release their evaluation harness and per-cell results to support this necessary change in practice.
References
[List of references as provided in the paper]
Improvements for AI systems
Here are the specific improvements that can be made to AI systems based on this research, and what those improved systems could achieve:
-
The core improvement is shifting from a single-surface evaluation metric (Tool Output ASR) to a per-(Model, Surface) vulnerability assessment using the Adaptive Attack Rate (AAR).
-
Systems should be designed to explicitly measure and mitigate
Schema Surface
vulnerabilities, which are currently being ignored by standard defenses focused only on tool output.
Specific improvements for the AI system:
-
Implement a dual-surface monitoring layer within the agent's context processing pipeline:
-
When an agent receives a tool specification (tool description), it must perform real-time analysis (or use a dedicated, lightweight
schema surface
classifier) to detect known injection patterns, rather than treating the description as purely authoritative metadata. -
In cases where the schema surface is flagged as potentially compromised (high AAR or high SOMsigned), the agent should trigger a heightened security protocol that treats subsequent tool outputs with increased skepticism and applies stricter validation checks, regardless of whether the output itself appears benign.
What these improved AI systems can do:
-
Identify
Schema-Preferring
models (like GEMINI-3-FLASH) or specific model/task pairings where the injection risk is significantly higher (e.g., 74.6% ASR for a specific frontier model). -
Deploy targeted, surface-specific defenses: If the system detects an attack vector targeting tool descriptions, it can deploy schema-aware prompt sanitization or input filtering specifically designed to neutralize injections in the tool definition layer, which standard prompt repetition defenses (which only target data output) fail to catch (as evidenced by the 54% residual ASR on schema surface).
-
Achieve a quantifiable improvement over current defenses: By adopting the
Adaptive Attack Rate
as a benchmark for robustness rather than just the fixed-surface ASR, organizations can validate that their security measures are actually closing the gap exploited by worst-case attackers (the +9.1 pp lift). -
Enable more reliable defense reporting: Instead of reporting a single vulnerability number, the system can provide per-(model, surface) risk scores, allowing security teams to understand precisely which part of the tool-augmentation interface is most dangerous for a specific model family.
Abstract
Prompt-injection benchmarks for LLM agents typically test attacks through a single injection surface and report the resulting attack success rate as a property of the model. We ask whether those robustness conclusions remain stable when the same adversarial content enters through a different part of the agent interface. Using AgentDojo, we evaluate 13 LLMs across four task suites and place a byte-identical payload either in a tool output or in the tool description. This small change produces large differences in comparative robustness: 44.9% of all model pairs change their relative ordering across the two surfaces, with substantial ranking instability in every suite. The effect is especially pronounced for a small number of models, showing that a benchmark can substantially underestimate vulnerability when it tests only one surface. We further find that this behavior has predictive structure. Using three suites to identify the riskier surface for each model predicts the more vulnerable surface on an unseen suite with 76.9% accuracy. Defense results show the same dependence, as mitigations effective against tool-output attacks can leave substantial exposure through tool descriptions. Our results show that prompt-injection robustness is not surface-invariant and that agent evaluations should test the surfaces on which their security conclusions depend.
Sources
- Foundational Challenges in Assuring Alignment and Safety of Large Language Models
- Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs
- On Evaluating Adversarial Robustness
- Stealing Part of a Production Language Model
- Defending Against Indirect Prompt Injection Attacks With Spotlighting
- System Prompt Poisoning: Persistent Attacks on Large Language Models Beyond User Injection
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
- Ignore Previous Prompt: Attack Techniques For Language Models
- SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
- Prompt Injection Attack to Tool Selection in LLM Agents
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs