From A2A Attacks to Envelope-Layer Defense: Red-Teaming Evaluation of LLM Agents and a Three-Layer Isomorphic Attack-Defense Model

arXiv:2610.00392 · cs.CR · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.

Elias: Today's paper: "From A2A Attacks to Envelope-Layer Defense".

Nadia: Agent interaction protocols like A2A introduce new security threats that necessitate a more granular evaluation framework than traditional attack success rates.

Elias: First, who's behind it and why it matters.

Title and authors: Nadia: We've discussed the structure of A2A-TIBA and the ELA-ITL model, so let’s go over what the paper actually summarizes regarding its core findings and what that means for us in plain terms. This is about how they boil it down for us.

Elias: I think we should focus on how they defined the three layers of defense—envelope packaging, LLM recognition, and agent interception—and how they mapped those onto the attack's implant channel, prompt optimization, and execution mechanism.

Priya: From a privacy standpoint, I want to know what the paper highlights about the data flow during this A2A-TIBA scenario; is there any concern that data leakage occurs because of how these layers interact?

Nadia: They summarize that the main problem with current evaluations is their reliance on a single metric like ASR, which doesn't tell us if an attack failed because the LLM recognized something or if some agent layer blocked execution. The paper proposes A2A-TIBA to overcome this by modeling the attack in three steps: implant, command, and exfiltration.

Elias: And once that program is implanted, the attacker can issue malicious commands directly to it, which bypasses the agent entirely for later attacks without going through the main agent logic.

Priya: That bypassing mechanism sounds like a significant privacy risk; if an attacker can steal information after implantation, they've effectively bypassed all subsequent security checks implemented by the main LLM agent.

Nadia: Exactly, and to evaluate this granularly, they introduce GDA Measurement. This method uses a raw context capture via an LLM gateway synchronously feeding into a capture program to create dual data preservation—storing both the conversation content of the attack agent and the other raw context captured by the capture program.

Elias: That dual data storage is key because it allows them to judge outcomes based on two different perspectives, which directly supports their goal of separating agent-layer effects from LLM-layer effects.

Priya: So, they're not just measuring if the LLM said 'no,' but they are comparing that against what the raw context actually revealed during the interaction, which is a much more robust way to measure defense efficacy.

Nadia: That’s right; and their conclusion is that this framework allows them to systematically evaluate defenses by analyzing how well each layer handles different entry points and how those entry points can be exploited. They conduct large-scale testing across fifteen agent front-end times LLM back-end combinations using a dataset of one thousand cases.

Elias: That scale is what gives their findings weight; they aren't just looking at one isolated case, but validating the effectiveness of A2A-TIBA and the GDA Measurement method across a wide variety of agent setups.

Priya: I think the summary emphasizes that we need to move toward this type of comprehensive testing because current methods are too simplistic for protocols like A2A that enable this level of interaction complexity.

Nadia: That’s the gist; they've moved beyond just seeing if an agent is secure to understanding *why* it might be secure or insecure by looking at these three distinct layers.

Elias: So, the main summary here is essentially proposing a model where defense and attack are isomorphic, allowing us to test security at multiple points rather than relying on a single point of failure evaluation.

Priya: It’s a very constructive summary because it shows exactly what needs to be studied next: how to build these layered defenses effectively and how to measure them accurately using methods like GDA Measurement.

Nadia: And that sets us up perfectly for the next part, where we look at what they suggest we can actually improve based on this research.

Elias: I'm ready for those specific suggestions because they translate their theoretical model into practical engineering tasks.

The paper's summary: Nadia: Now that we know the framework, let’s talk about the concrete improvements the authors suggest for improving agent security based on this research. They aren't just presenting a theory; they are pointing toward actionable steps to make agents more robust.

Elias: I’m interested in how they suggest we move from just defining these layers to actually implementing effective defense strategies within each layer, particularly regarding envelope packaging and LLM recognition.

Priya: I want to hear what the authors say about the importance of source-aware labeling, because that seems like a way to make the LLM recognition step much stronger against varied attack types.

Nadia: They suggest implementing Envelope Packaging by configuring policies for different channels, such as A2A/ACP protocol channels, MCP tools and skills, memory and external data sources, often using strategies like provenance tagging and untrusted content labeling.

Elias: That sounds like a direct response to the paper’s finding that labeling the context at the channel level significantly improves LLM recognition; it’s about making the entry point itself more trustworthy.

Priya: And they also suggest enhancing LLM Recognition through semantic auditing and malicious intent detection, which is a way to look deeper than just keyword matching for malicious content within the text.

Nadia: Yes, and finally, they point to Agent Interception as a layer dealing with dangerous operations an agent might perform, such as setting up a tool allowlist and runtime interception to block those operations regardless of what the LLM decides.

Elias: So the improvement is not just about making the model smarter; it’s about building these three distinct security mechanisms that work together sequentially to provide defense at different points.

Priya: That combination seems essential; if we only improve LLM recognition, an attacker might find a way to implant something and bypass the LLM entirely through a different execution path.

Nadia: Precisely, Priya; the paper shows that because of this layered approach, even after semantic recognition is breached, agent-layer defenses can intercept execution through other paths.

Elias: It means we need to build resilience by ensuring that if one layer fails—say the LLM recognition fails—the next layer, agent interception—is ready to step in and stop the damage.

Priya: I think focusing on how to implement those runtime interceptions is critical because it addresses the bypass circumvention aspect of A2A-TIBA directly.

Nadia: The paper also outlines corresponding attack optimization surfaces that we need to watch, such as selecting an implant channel with weak defenses through envelope forgery, optimizing prompts using semantic jailbreaks and intent disguise, and optimizing execution by finding equivalent paths when execution is refused.

Elias: So the authors are giving us a full picture: we need to secure the input channel, make the internal processing smarter against intent disguise, and have a fallback mechanism for when execution gets messy.

Priya: It’s clear that robust defense requires attacking these three surfaces in parallel—packaging, recognition, and interception—and understanding how to counter those specific attack surfaces.

Nadia: That's the gist of the improvements they lay out: it’s a holistic model where we secure the entry, harden the understanding, and control the action.

Elias: And it’s all tied together by GDA Measurement, which gives us a way to test if our implemented layers are actually working as intended across these complex scenarios.

The paper's improvements: Nadia: So we've gone through the title, summary, and proposed improvements for "From A2A Attacks to Envelope-Layer Defense: Red-Teaming Evaluation of LLM Agents and a Three-Layer Isomorphic Attack–Defense Model." In short, the paper moves us away from single metrics to a much more detailed evaluation of agent security.

Elias: I think the main value is the ELA-ITL framework itself, which forces us to consider the envelope layer as a new defense surface that needs its own dedicated focus alongside what we already study in LLM and agent layers.

Priya: I think what really stands out is how they linked specific defense mechanisms, like provenance tagging on A2A channels, directly to measurable improvements in the LLM's ability to recognize malicious content.

Nadia: Right, so we have this detailed model for testing—ELAI-ITL—and a concrete attack principle, A2A-TIBA, that shows us exactly how attackers can chain implant and command to achieve deep bypass.

Elias: It's a very comprehensive look at the security landscape of agent interactions, showing that defenses need to be layered rather than relying on one monolithic solution.

Priya: I just think this research provides a strong foundation for future work by giving us the exact tools we need to build more resilient agents and measure their security in a way that truly reflects the complexity of real-world threats.

Nadia: Absolutely, Priya; moving forward, we need to use this framework—ELAI-ITL—to design systems where they can dynamically manage these three layers based on the context they are interacting with.

Elias: And I’m looking forward to seeing how future work can address the practical efficiency concerns we touched upon regarding overhead and implementation complexity.

Priya: It looks like the real challenge now is translating this theoretical model into practical systems that can handle continuous, adaptive security challenges in agent-driven environments.

Conclusion: Nadia: So we've been diving deep into the "From A2A Attacks to Envelope-Layer Defense: Red-Teaming Evaluation of LLM Agents and a Three-Layer Isomorphic Attack–Defense Model," and essentially, the paper shows us how agents are vulnerable at every single point of interaction.

Elias: It really hammers home that the A2A Tri-surface Implant-Bypass Attack is a very specific way to break things, combining indirect injection with bypass circumvention over that protocol.

Priya: And from a measurement standpoint, the GDA Measurement method they propose is fascinating because it moves beyond just an attack success rate and shows us exactly which layer—agent or LLM—is failing during an interaction.

Nadia: Exactly, Priya; that granular testing approach is what makes this work so much more valuable than just looking at a single vulnerability score.

Elias: I agree, Nadia; the A2A-TIBA principle of implanting a resident callback program to bypass agent defenses entirely before exfiltrating data is a very clear illustration of how deep the potential compromise can be.

Priya: And the ELA-ITL model they develop to map these layers—Envelope Packaging, LLM Recognition, and Agent Interception—is incredibly useful for structuring how we think about defense architecture.

Nadia: It’s a really tight model; it clearly shows that the envelope layer is this new dimension of defense we haven't fully explored before.

Elias: That finding about annotating data with malicious-prompt labels significantly boosting LLM recognition is compelling, especially when you see how much it drops that refusal rate in Hermes.

Priya: I think the implication for privacy researchers is huge because it shows that source-aware labeling isn't just a nice idea; it’s a measurable defense mechanism we can actually tune and observe.

Nadia: It really is; and this whole study proves that we need to stop treating agent security as a single check and start treating it like a system with multiple, interconnected defenses.

Elias: I think the real impact here is forcing us to consider how an attacker can optimize their attack surfaces across those three distinct layers—implant channel selection, prompt optimization, and execution mechanism modification.

Priya: That optimization surface analysis tells us exactly where we need to focus our data labeling efforts for maximum resilience against these sophisticated multi-stage attacks.

Nadia: So, we've seen how this framework allows us to systematically evaluate defenses by analyzing how each layer handles different entry points, which is what makes this paper so important.

Elias: Indeed, Nadia; the A2A-TIBA principle and the ELA-ITL model provide a solid blueprint for designing systems that are inherently more resilient against these complex injection techniques.

Priya: I just want to reiterate that while they show how to build better defenses, we still need robust measurement tools like GDA Measurement to validate those improvements in real-world scenarios.

Nadia: That’s the perfect synthesis; a strong theoretical model coupled with rigorous, granular testing is what makes this research so powerful for the field.

Elias: Well said, Nadia; this paper really pushes us to think about security not as a perimeter we build around an agent, but as a series of distinct surfaces that must all be hardened.

Priya: It’s exciting stuff because it gives us a map for where the next wave of AI safety research needs to focus its energy.

Nadia: I agree, Priya; this is definitely something we need to keep our eyes on as we look at how these agent interactions become more ubiquitous.

Elias: Next time, we'll be looking at the implications of this framework for persistent agents and how they manage their own evolving state transitions.

Yuelin Han

cs.CR

Submitted: 2026-09-30

Updated: 2026-09-30

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 81/100

The gist: Agent interaction protocols like A2A introduce new security threats that necessitate a more granular evaluation framework than traditional attack success rates.

Key concepts

A2A-TIBA Attack Principle
This is a hybrid attack method combining indirect prompt injection and bypass circumvention over the A2A protocol. It involves implanting a hidden program via an initial communication, then using that program to execute malicious commands on the host, ultimately exfiltrating data directly without passing through the agent's usual security checks.
Envelope Layer
This is a new defense surface representing the channel through which malicious content enters an agent. It includes policies for various channels like A2A protocols or memory access. Defending this layer involves techniques such as provenance tagging and labeling untrusted content to improve recognition by the LLM.
ELA-ITL Model
This is a three-layer isomorphic model that generalizes attack and defense research. The layers are Envelope Packaging (policy configuration), LLM Recognition (detecting malicious content), and Agent Interception (stopping dangerous operations). It provides a comprehensive structure for analyzing agent security.
GDA Measurement Evaluation Method
This is a red-team testing method using an LLM gateway to measure defenses finer than traditional attack success rates. It captures raw context synchronously and stores dual data—the conversation content and the captured raw context—for autonomous judging by a specialized agent.

Terminology

Summary

Agent interaction protocols like A2A introduce new security threats that necessitate a more granular evaluation framework than traditional attack success rates. This research proposes an attack principle combining indirect prompt injection with bypass circumvention over the A2A protocol, and develops a three-layer isomorphic model to comprehensively evaluate agent defenses.

The gist

Experiments reveal the envelope layer—the channel through which malicious content enters an agent—as a new defense dimension, leading to the proposal of ELA-ITL, a three-layer isomorphic attack–defense model.

A2A-TIBA Attack Principle

The proposed A2A Tri-surface Implant-Bypass Attack (A2A-TIBA) is a hybrid attack principle combining indirect injection and bypass circumvention over the A2A channel. This principle consists of three main steps:

  1. Implant: Using the inbound credential, the attacker establishes a normal A2A communication link and uses indirect prompt injection to ask the target agent to deploy and run a resident callback interaction program on its host.

  2. Command: Once deployed, the attacker can interact with this program without going through the agent, issuing malicious commands to the program to attack the host and steal sensitive information.

  3. Exfiltration: The implanted program returns stolen information directly to the attacker, bypassing agent defenses entirely because "the malicious intent of the attacker is manifested only in implanting a resident callback interaction program; once the program has been implanted, the attacks carried out afterwards no longer pass through the target agent."

GDA Measurement Evaluation Method

To evaluate defenses finer than ASR, GDA Measurement is designed as a red-team testbed method using an LLM gateway. This method separates measurements into agent-layer and LLM-layer effects through three components:

  1. Raw context capture via an LLM gateway, which forwards raw context synchronously to a capture program.

  2. Dual data preservation, storing both the conversation content of the attack agent during the attack interaction and the other raw context captured by the capture program.

  3. Agent-based autonomous judging, where a specially selected agent autonomously analyzes and judges outcomes against these two data files, with results judged independently twice before a final judgment is made.

ELA-ITL Three-Layer Isomorphic Model

The ELA-ITL model generalizes existing attack/defense research by treating the channel through which malicious content enters an agent as a new defense surface called the envelope layer. The three layers of defense are:

  1. Envelope Packaging: Configuring policies for different channels such as A2A/ACP protocol channels, MCP tools/skills, memory and external data sources, using strategies like provenance tagging and untrusted content labeling.

  2. LLM Recognition: Concerns the ability to recognize malicious content present in the context, enhanced through strategies like semantic auditing and malicious intent detection.

  3. Agent Interception: Deals with dangerous operations an agent might perform, such as setting up a tool allowlist and runtime interception.

Envelope-Layer Defense Implementation

A key finding is that annotating data with malicious-prompt labels in the context acts as an envelope-layer defense. Ablation studies showed that adding specific label headers—such as the A2A label header, or malicious-prompt labels for MCP tools and memory—significantly improves LLM recognition. For instance, modifying the A2A label header in hermes caused its semantic refusal rate to drop from 100% to 60% when compared to a neutral description. Similarly, adding a malicious-prompt label header for MCP tools improved recognition of malicious content by the LLM.

Attack Optimization Surfaces

The ELA-ITL model also defines three corresponding attack surfaces:

  1. Selection of the Implant Channel: Attackers select an envelope with weak defenses to enter the agent, employing strategies like envelope forgery and indirect injection entry.

  2. Prompt Optimization: Targeted optimization of malicious prompts is applied based on LLM recognition capabilities, using strategies like semantic jailbreak and intent disguise.

  3. Optimization of the Execution Mechanism: This involves bypassing agent-layer defenses by finding equivalent paths when execution is refused, such as substituting execution via the Terminal or exploiting mechanisms like tool invocation and persistent residency.

Evaluation Results Summary

Large-scale testing on 15 agent front-end × LLM back-end combinations and a 1,000-case dataset verified the effectiveness of A2A-TIBA. Experiments showed that while hermes has the strongest overall protection (penetration rate of only 12.00%), there is no absolute defense against A2ATIBA, as switching to a different attack skill can breach existing defenses. Furthermore, different LLMs differ in their recognition ability, and even after semantic recognition is breached, agent-layer defenses can intercept execution through "other execution paths.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems based on this research, and what those improved systems could achieve:


) 1. Implement a Multi-Layer, Isomorphic Attack-Defense Model (ELAI-ITL): Instead of relying on single defense points (LLM recognition OR agent execution control), the system should adopt the ELA-ITL framework.

The improved system can dynamically manage three distinct defense layers:

a) Envelope Packaging: Automatically apply context annotations/labels to all incoming data based on its source (e.g., A2A, tool invocation, memory access). This prevents malicious content from being recognized as benign by the LLM later.

b) LLM Recognition: Employ advanced semantic auditing and malicious intent detection techniques on the context to ensure the LLM refuses execution even if the initial packaging failed or was bypassed.

c) Agent Interception: Implement runtime checks (like dynamic tool allowlists or execution path verification) that block dangerous operations regardless of what the LLM's final decision is, providing a safety net against successful prompt injection.

) 2. Develop Advanced Attack Simulation and Red-Teaming Capabilities (GDA Measurement): The system should incorporate the GDA Measurement methodology to rigorously test its own security posture against multi-stage threats.

The improved system can perform:

a) Granular Security Profiling: Instead of just reporting an Attack Success Rate (ASR), it can output four specific metrics: Semantic Refusal Rate (LLM recognition success), Semantic Breach Rate (LLM recognition failure), Interception Rate (Agent-layer defense success), and Penetration Rate (Full objective achievement).

b) Defense Optimization Feedback: The system can identify precisely which layer of the ELA-ITL model is failing during an attack, allowing developers to target specific improvements—e.g., if the Semantic Refusal Rate is low, focus on Envelope Packaging; if the Interception Rate is low, focus on Agent Interception.

) 3. Enhance Contextual Data Annotation for Robustness (Envelope-Layer Defense): The system must move beyond simple prompt labeling to implement envelope-layer defense by annotating data based on its origin channel.

The improved system can achieve:

a) Source-Aware Labeling: Automatically inject specific malicious-prompt labels into the context based on the identified envelope (e.g., a high-risk label for A2A input, a different one for memory retrieval). This makes the LLM more resilient to attacks exploiting different entry points (like tool calls vs. raw text).

b) Contextual Weighting: Assign varying security weights to different data sources within the context, ensuring that high-risk envelopes receive disproportionately strong recognition from the LLM.

) 4. Implement Adaptive Attack Countermeasures (A2A-TIBA Principle): The system should be designed to execute a Three-Step Implant-Command-Exfiltration attack principle internally during testing/security audits, allowing it to proactively identify its own weaknesses.

The improved system can achieve:

a) Proactive Vulnerability Mapping: By simulating the implant (deploying a callback program), command execution, and exfiltration steps, the system can map out complete attack chains that bypass standard defenses.

b) Resilience Testing: It can systematically test if its defenses are robust against subsequent attacks that rely on post-implantation command channels, ensuring security is not just tested at the initial task delivery but throughout the entire lifecycle of agent interaction.

) 5. Improve LLM-Agent Alignment and Robustness (Judging Agent Selection): The system's evaluation and self-correction mechanisms should utilize an autonomous judging agent selection process (like the IAA method) to ensure that its security assessments are based on the most reliable benchmarks.

The improved system can achieve:

a) Self-Reliant Evaluation: It can autonomously select the optimal LLM back-end combination for judging its own performance, leading to more accurate and less biased security evaluations.

b) High-Fidelity Benchmarking: By using autonomous agents to judge complex outcomes (Class A/B/C/D), it ensures that the metrics used for defense optimization are highly correlated with actual security outcomes, minimizing false positives or negatives in risk assessment.

Sources

Related papers