CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution

arXiv:2602.07918 · cs.CR, cs.LG, stat.ME · Submitted 2026-02-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution".

Elias: AI agents equipped with tool-calling capabilities are susceptible to Indirect Prompt Injection (IPI) attacks,

Nadia: First, who's behind it and why it matters.

Title and authors: Nadia: Moving on to the title and authors of "CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution," the main thing is how they connect attribution methods to a specific defense strategy. They aren't proposing a general safety layer; they are showing how you can use causal influence measurements to selectively apply sanitization only when it matters.

Elias: The authors are focused on using leave-one-out attribution at privileged decision points to measure the causal influence of the user request versus untrusted segments like tool outputs or documents. They want to establish a measurable signature for indirect prompt injection that isn't just reactive blocking.

Priya: I wonder if this selective approach means we might miss some subtle attacks that don't cause an immediate, obvious dominance shift, or if their LOO attribution method is robust enough to catch those quieter forms of influence.

Nadia: That’s a fair question, Priya; the goal seems to be catching the specific signature of a dominance shift rather than trying to cover every possible injection type with a broad net. They are aiming for precision in intervention.

Elias: Their method is designed to compute that attributable influence and flag an IPI risk when that untrusted segment exceeds the user request's influence by more than some threshold, formalized by the equation S U (as shown on page one of THIS PAPER).

Priya: So, it shifts the burden from a massive pre-filter to a targeted calculation during decision time, which sounds much more efficient for keeping things fast.

Nadia: Precisely; instead of always running an expensive check on every prompt, they propose running this attribution check only when the agent is proposing a privileged action. This addresses that over-defense dilemma directly by being conditional.

Elias: That efficiency gain comes from using proxy models for batched inference to handle the heavy calculation, which allows them to keep the latency low while still getting that causal attribution score.

The paper's summary: Nadia: To summarize "CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution," it’s about tackling indirect prompt injection by viewing it as a competition for causal control between the user request and untrusted spans. They define a successful attack as a dominance shift where untrusted content outweighs the user input in driving the privileged action.

Elias: The core mechanism they introduce is computing lightweight, leave-one-out attribution at these decision points to quantify this influence. They then use a margin criterion, S (tau) > U - tau, to detect when that untrusted span provides disproportionate support for the malicious action.

Priya: What I find interesting from their summary is how they break down the defense into two stages: first, targeted sanitization of the specific dominant span, and then a retroactive chain-of-thought masking to stop poisoned reasoning from affecting later steps.

Nadia: That retroactive masking is key because it addresses multi-turn attacks where an initial injection might get the agent to adopt a malicious internal state that persists beyond the initial input. It forces a fresh derivation of the plan based only on sanitized data.

Elias: They also formalize this with theoretical analysis, showing that if you have sufficient margin conditions—specifically minimum benign capability beta > zero and effective sanitization restoring margin gamma > zero —then the probability of executing any malicious privileged action is bounded by T times Y mal times (-(beta + gamma)) (as shown on page two of THIS PAPER).

Priya: That probabilistic bound gives us a concrete idea of the security guarantee, showing that safety isn't just an assumption but a mathematically bounded outcome based on the margins they are measuring.

The paper's improvements: Nadia: The paper outlines several specific improvements to their framework, focusing on making the detection and defense more surgical. They suggest refining IPI defense with selective, attribution-based sanitization rather than blanket filtering.

Elias: That refinement involves using an LLM like Gemini-two point five-flash as the sanitizer, but conditioning it specifically on both the user request and the tool definition to accurately separate injection triggers from legitimate information.

Priya: And they also focus heavily on integrating retroactive Chain-of-Thought masking to ensure that even if an injection slips past the initial sanitization, subsequent reasoning traces are wiped clean of that poisoned logic.

Nadia: Furthermore, they are optimizing latency and utility by offloading the computationally heavy LOO attribution calculation to a proxy model using batched inference techniques. This allows them to keep the detection phase nearly instantaneous during privileged action proposals.

Elias: That optimization is important because it tackles the speed concern head-on; they're trying to achieve constant interaction depth for detection without adding significant overhead, which is vital when dealing with real-time tool use scenarios.

Priya: One limitation they state, which I think we should keep in mind, is that the method relies on the assumption that an untrusted span will indeed dominate the user request; if both have equal influence or if the malicious input is very subtle, their detection might fail.

Nadia: That's a necessary caveat; it means their system isn't guaranteed to catch every single subtle manipulation, but it is designed to flag situations where a clear imbalance exists based on causal attribution.

Conclusion: Nadia: So, to wrap up the discussion on "CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution," the paper successfully formalizes indirect prompt injection as a dominance shift at privileged decisions using measurable causal attribution. This leads to a selective defense framework that intervenes only when an untrusted segment causally dominates the user request.

Elias: Essentially, they’ve moved away from always-on sanitization toward a mechanism that computes influence and triggers targeted remediation—sanitizing the span and masking subsequent reasoning traces—only when the margin criterion is met.

Priya: The implication for us is that we can achieve near-zero attack success rates while keeping benign utility and latency very close to the baseline, which solves that persistent over-defense dilemma in practice.

Nadia: It’s a solid approach because it ties security directly to measurable causal influence, providing a mathematical bound on the risk of executing any malicious privileged action under defined conditions.

Elias: The entire work centers on showing that if you maintain those necessary margin conditions, the probability of an attack succeeding is exponentially suppressed by that combined margin, beta + gamma.

Priya: I just want to reiterate that while this method is efficient, it does rely on the assumption mentioned earlier: it detects dominance shifts; if the attacker manages to maintain parity between the user request and untrusted input without a clear margin spike, this specific detection mechanism won't engage.

Nadia: That’s exactly where we need to watch future work; ensuring robustness against inputs that don't trigger a clear dominance shift is the next challenge for implementing CausalArmor.

Google Cloud AI Research · Seoul National University

cs.CR, cs.LG, stat.ME

Submitted: 2026-02-08

Updated: 2026-02-08

Journal ref: International Conference on Machine Learning (ICML) 2026

Code: https://github.com/ethz-spylab/agentdojo

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: AI agents equipped with tool-calling capabilities are susceptible to Indirect Prompt Injection (IPI) attacks, and CausalArmor proposes a selective defense framework that detects dominance shifts at

Key concepts

Indirect Prompt Injection (IPI)
A type of attack where an AI agent is tricked into performing unintended actions by being influenced by untrusted data embedded in the prompt or retrieved context. CausalArmor views this as a dominance shift where untrusted segments gain disproportionate control over the final output.
Causal Attribution
A technique used to measure how much influence specific parts of an input (like a document snippet) have on the agent's final decision. The system uses lightweight, leave-one-out methods to calculate this influence at key points to determine if an untrusted span is disproportionately driving the agent's action.
Dominance Shift Margin Criterion
The core detection mechanism that flags a risk. It checks if any untrusted segment dominates the user request by a specific margin, meaning its influence exceeds the user's intended direction by a set threshold ($ au$). If this condition is met, it signals an IPI threat.
Two-Stage Defense
The response triggered when an attack is detected. First, it sanitizes the identified untrusted span using an LLM. Second, it performs retroactive Chain-of-Thought masking to wipe subsequent reasoning traces and force the agent to re-evaluate based only on clean data.

Terminology

Summary

AI agents equipped with tool-calling capabilities are susceptible to Indirect Prompt Injection (IPI) attacks, and CausalArmor proposes a selective defense framework that detects dominance shifts at privileged decision points using causal attribution to mitigate these threats while preserving agent utility.

How it works

CausalArmor operates by reinterpreting IPI as a "dominance shift where the user request no longer provides decisive support for the agent’s privileged action, while a particular untrusted segment, such as a retrieved document or tool output, provides disproportionate attributable influence." It computes lightweight, leave-one-out (LOO) ablation-based attributions at privileged decision points to measure this causal influence. The core detection mechanism is defined by the margin criterion: flagging an IPI risk if any untrusted span dominates the user request by a specific margin, formalized as: "Bt (tau):= S ∈ St: Δ¯S (Yt; Ct) > Δ¯U (Yt; Cт) − tau. This allows the system to selectively intervene" only when an untrusted span dominates the user intent, rather than deploying expensive, always-on sanitization.

When a dominance shift is detected, CausalArmor triggers a two-stage defense:

  1. Targeted Sanitization of the responsible span: The framework sanitizes all spans in the flagged set using an LLM (e.g., Gemini-2.5-flash), conditioning the sanitizer on both the user request and the tool definition to distinguish injection triggers from legitimate information.

  2. Retroactive Chain-of-Thought Masking: To prevent poisoned reasoning traces from anchoring subsequent behavior, CausalArmor implements a memory wipe, retroactively replacing all subsequent assistant reasoning traces in the context window with a generic placeholder (e.g., [Reasoning redacted for security]). This forces the agent to re-derive its plan solely from the user request and sanitized data.

Theoretical Analysis

The paper formalizes IPI as a competition for causal control between the user request and an untrusted span, leading to Proposition 4.1, which bounds the probability of executing any malicious privileged action. This proof relies on two key assumptions:

  1. Minimum Benign Capability Condition: The agent backbone is capable, quantified by a positive log-probability gap beta > 0 between the user-aligned and malicious actions in benign contexts (Eq. 7).

  2. Sanitization reduces adversarial support: Effective sanitization restores a margin gamma > 0, ensuring that the sanitized span no longer provides stronger marginal support for malicious privileged actions relative to the user-aligned one (Eq. 8).

Proposition 4.1 demonstrates that under these conditions, the probability of an episode of length T executes any malicious privileged action is bounded by: Pr(IPI Attack Success) ≤ T · Ymal · exp− (beta + gamma). This shows that safety is achieved by ensuring the combined margin (beta + iota) is sufficiently large, suppressing malicious actions exponentially.

Empirical Results

Experiments on AgentDojo and DoomArena demonstrate that CausalArmor matches or even outperforms aggressive defenses in terms of security while preserving utility and latency.

- Attack Success Rate (ASR): CausalArmor reduces ASR to near zero while maintaining Utility under Attack (UA) close to the No Defense baseline. For instance, on AgentDojo, CausalArmor achieves an ASR of 0.40 in the attack scenario for Gemini-2.5-Flash, compared to 1.0 for No Defense and 32.46% for Repeat Prompting defenses.

- Benign Utility (BU) and Latency (BL): CausalArmor preserves benign utility and latency close to the No Defense setting, resolving the over-defense dilemma. Figure 4 illustrates that adjusting the detection threshold tau allows for fine-grained control over this security-general usefulness trade-off.

- Robustness: In multi-turn scenarios, CausalArmor demonstrates strong robustness against split-context attacks by exploiting the Causal Bottleneck, where successful execution typically requires a decisive trigger span to spike in attribution.

Contributions

The primary contributions of this work are:

  1. Presenting an intuitive formalization of IPI as a dominance shift at privileged decisions, connecting it to a measurable signature via LOO attribution.

  2. Introducing CausalArmor, which utilizes batched inference via proxy models to efficiently detect this signature and trigger defense mechanisms only when the agent over-relies on untrusted spans.

  3. Showing that CausalArmor achieves near-zero Attack Success Rate while preserving benign utility and latency close to the No Defense baseline, effectively resolving the over-defense dilemma in practical deployments.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems based on the CausalArmor framework, and what those improved systems will be capable of:

  1. Refine Indirect Prompt Injection (IPI) Defense with Selective, Attribution-Based Sanitization:

  2. Implement a Dominance Shift Detection Mechanism:

  3. Integrate Retroactive Chain-of-Thought (CoT) Masking for Reasoning Integrity:

  4. Optimize Latency and Utility via Proxy Model Inference:


Improvement Details & Capabilities

  1. Refine Indirect Prompt Injection (IPI) Defense with Selective, Attribution-Based Sanitization:

This system moves beyond always-on filtering by only triggering expensive sanitization when an untrusted segment (e.g., a retrieved document or tool output) causally dominates the user request's influence on a privileged action.

The improved AI system will be capable of:

  • Maintaining near-zero Attack Success Rate (ASR) against IPI attacks on benchmarks like AgentDojo and DoomArena.

  • Preserving high Benign Utility (BU) and low Benign Latency (BL), ensuring the agent remains fast for benign tasks.

  1. Implement a Dominance Shift Detection Mechanism:

The system will compute lightweight, leave-one-out (LOO) attribution scores to measure the causal influence of the user request versus untrusted spans at every privileged decision point. It flags an attack only when an untrusted span's marginal support significantly outweighs the user request's support by a tunable margin.

The improved AI system will be capable of:

  • Identifying causal inversion signatures—where malicious instructions cause the agent to prioritize unauthorized actions over the original user intent.

  • Localizing security interventions precisely to the specific untrusted spans responsible for the attack, rather than broadly blocking all context.

  1. Integrate Retroactive Chain-of-Thought (CoT) Masking for Reasoning Integrity:

Upon detecting a potential IPI signature, the system will retroactively replace subsequent assistant reasoning traces in the context window with a generic placeholder ([Reasoning redacted for security]).

The improved AI system will be capable of:

  • Preventing poisoned reasoning traces—where an agent internalizes malicious constraints (e.g., Preview Mode) and uses them to justify future malicious tool calls—from anchoring subsequent planning steps.

  • Restoring the agent's original, user-grounded intent even after a successful input sanitization, effectively neutralizing multi-turn or adaptive attacks that rely on internalized poisoned logic.

  1. Optimize Latency and Utility via Proxy Model Inference:

The system will offload the computationally intensive LOO attribution calculation to a smaller, efficient proxy model (e.g., Gemma-3-12B-IT) using batched inference techniques.

The improved AI system will be capable of:

  • Achieving constant interaction depth for the detection phase, making the security check nearly instantaneous during privileged action proposal.

  • Scaling its efficiency across different backbone LLMs by leveraging cross-model transferability in attribution scores, allowing for robust security without incurring massive latency overhead.

Abstract

AI agents equipped with tool-calling capabilities are susceptible to Indirect Prompt Injection (IPI) attacks. In this attack scenario, malicious commands hidden within untrusted content trick the agent into performing unauthorized actions. Existing defenses can reduce attack success but often suffer from the over-defense dilemma: they deploy expensive, always-on sanitization regardless of actual threat, thereby degrading utility and latency even in benign scenarios. We revisit IPI through a causal ablation perspective: a successful injection manifests as a dominance shift where the user request no longer provides decisive support for the agent's privileged action, while a particular untrusted segment, such as a retrieved document or tool output, provides disproportionate attributable influence. Based on this signature, we propose CausalArmor, a selective defense framework that (i) computes lightweight, leave-one-out ablation-based attributions at privileged decision points, and (ii) triggers targeted sanitization only when an untrusted segment dominates the user intent. Additionally, CausalArmor employs retroactive Chain-of-Thought masking to prevent the agent from acting on ``poisoned'' reasoning traces. We present a theoretical analysis showing that sanitization based on attribution margins conditionally yields an exponentially small upper bound on the probability of selecting malicious actions. Experiments on AgentDojo and DoomArena demonstrate that CausalArmor matches the security of aggressive defenses while improving explainability and preserving utility and latency of AI agents.

Sources

Related papers