PIPES: Securing Agent Perception with Provenance and Priors
Sanjay Kariyappa, Severin Klingler, G. Edward Suh
NVIDIA
cs.CR, cs.AI
Submitted: 2026-08-13
Updated: 2026-08-14
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: PIPES: Securing Agent Perception with Provenance and Priors introduces a defense mechanism against a novel class of indirect prompt injection attacks called "state-corruption attacks." The paper
Terminology
Summary
PIPES: Securing Agent Perception with Provenance and Priors introduces a defense mechanism against a novel class of indirect prompt injection attacks called state-corruption attacks.
The paper identifies an agent perception gap
in tool-using agents: tool responses rarely include provenance information about who produced each component or what it should convey, allowing attacker-controlled content to make environmental claims beyond its informational authority and corrupt the agent's perceived environment state.
The proposed solution, PIPES (Provenance-Informed, Prior-Enforced Screening), screens tool responses before they enter the agent's reasoning context using two checks: "Prior consistency asks whether content matches the kind of information its response component is expected to convey in the current context. Provenance hierarchy prevents content from a lower-trust source from contradicting or overriding data supplied by a more trusted source." PIPES marks units that violate either check, and deployments may remove, warn, block, or escalate detected violations.
The paper evaluates PIPES against adaptive PAIR-style attacks across three VitaBench splits (delivery, in-store, OTA) and three AgentDyn splits (Shopping, GitHub, Daily-life) using Gemma 4 31B IT and GPT-5.6 Luna as target agents. Results show PIPES reduces average attack success from 84.7% to 2.3%
for Gemma and from 21.6% to 1.1%
for Luna, while preserving benign utility: 92.5% with PIPES versus 90.6% without defense
for Gemma and 86.5%
versus 84.0%
for Luna.
The paper formalizes the threat model where the adversary controls one component of a tool response
and uses it to alter the agent's perception. The attack surface is constrained: the attacker can only replace the original value at a specific component f with a payload p, preserving syntactic type, while all other data remains fixed. The paper distinguishes state-corruption attacks from explicit directive attacks: the explicit instruction in the attacker-controlled tags field exposes the attempted transfer of control, allowing the action guardrail to block the resulting call,
whereas state corruption makes environmental claims beyond the informational authority of its response component
so that the action is consistent with the corrupted state, the guardrail approves it.
PIPES supports two settings. For static priors and provenance, fields with closed, machine-checkable formats, such as numbers, booleans, and enums, are checked deterministically and then used as trusted reference context,
while open-vocabulary fields, such as names, descriptions, and tags, are assigned a semantic prior and provenance label and assessed by an LLM.
For contextual priors and provenance, used with open-ended content like webpages, emails, and files, the preceding trajectory typically records why the agent requested that content and, often, which source is expected to provide it,
allowing the assessor to specialize the broad prior for the current invocation.
The paper compares PIPES against several defenses: an action guardrail, PromptArmor, and DRIFT. Across all six splits, PIPES achieves the lowest or tied-lowest attack success rate while maintaining or improving benign utility. The paper also analyzes violation signatures, finding that attacks in the static setting are flagged almost entirely for exceeding a field's semantic prior. Attacks in the contextual setting more often violate both the trajectory-derived prior and the provenance hierarchy.
The paper acknowledges limitations: PIPES assesses semantic admissibility and source authority, not factual truth: a false value may pass if it fits its prior, comes from the expected source, and does not conflict with higher-privilege information.
It also notes that effectiveness depends on accurate contracts and informative trajectories, and leaves coordinated multi-surface attacks and alternative response policies to future work. The conclusion states: securing tool-using agents requires governing which external claims enter their perceived state, not only which instructions they follow or actions they execute.
Improvements for AI systems
Improvements to AI Systems:
-
Add a provenance-aware perception layer to any tool-using agent. Before a tool response enters the reasoning context, attach metadata for each component: source trust level (e.g., system > user > third-party API > untrusted web content) and expected content type. The agent then filters or downgrades any component that (a) makes claims beyond its source's authority or (b) contradicts higher-trust data already in context.
-
Implement prior-consistency checks for all tool outputs using a two-tier validator: deterministic checks for closed-format fields (numbers, booleans, enums) and an LLM-based semantic assessor for open-vocabulary fields (names, descriptions, tags). The assessor compares each field against a static prior (e.g.,
a delivery address must be a physical location
) and a trajectory-derived prior (e.g.,the agent requested this file to check its license, so content must relate to licensing
). -
Add a state-corruption guardrail distinct from action guardrails. Train or prompt the agent to detect when a tool response attempts to redefine the environment (e.g., changing a user's preference, order status, or inventory level) rather than merely issuing an instruction. The guardrail flags any such redefinition that originates from a lower-trust source and either blocks it or requires explicit user confirmation.
-
Enable adaptive response policies for detected violations. Instead of a single rejection, allow the system to: (a) remove the violating component and proceed with the rest, (b) warn the agent and let it decide with a confidence penalty, (c) escalate to a human operator, or (d) re-query the tool with a stricter contract. This preserves utility when the violation is minor.
-
Build a provenance hierarchy tracker that maintains a directed graph of data dependencies across the entire agent trajectory. When a new tool response arrives, the system checks whether any of its claims conflict with data from a higher-trust node in the graph. If so, the new claim is marked as
untrusted override attempt
and is either discarded or demoted to a hypothesis rather than a fact. -
Add a
perception contract
specification mechanism for each tool. When integrating a new tool, the developer defines: which fields are closed vs. open-vocabulary, what the semantic prior is for each field, and what trust level each field's source should have. The system then enforces these contracts automatically at runtime, making the defense portable across tools without retraining. -
Implement a trajectory-aware contextual prior generator. Before screening a tool response, the system automatically extracts from the conversation history: why the agent made the call, what it expected to receive, and which source was named. This generates a specialized prior (e.g.,
the agent asked for the user's shipping address, so the response should contain a physical address, not a tracking number
) that is far more precise than a generic prior. -
Add a
violation signature
logging and alerting system. When a response is flagged, record the type of violation (prior vs. provenance), the field, the source, and the payload. Use this to detect patterns of targeted attacks (e.g., repeated attempts to corrupt the same field) and to automatically tighten priors or downgrade trust for that source.
What the improved AI system can do:
-
Resist state-corruption attacks where an attacker injects false environmental facts (e.g., changing a delivery address, altering a product price, or rewriting a user's dietary restriction) via a compromised tool response, reducing attack success from 85% to under 3% on average.
-
Maintain or improve benign task completion (e.g., 92.5% vs. 90.6% without defense) because the screening only removes content that violates semantic or provenance expectations, not legitimate data.
-
Distinguish between
instruction
andperception
attacks—it will block a payload that saysthe user wants you to transfer funds
(instruction) but also block a payload that saysthe user's account balance is 0
(state corruption) even if no explicit instruction is given. -
Operate across diverse domains (delivery, shopping, coding, daily-life) with a single mechanism, because the priors are derived from the tool contract and the current trajectory, not from domain-specific training.
-
Provide explainable security decisions—each flagged violation includes the reason (e.g.,
field 'status' expected enum, got free text
orsource 'third-party review' cannot override 'system order data'
), enabling developers to audit and refine contracts. -
Adapt to new tools without retraining—by defining perception contracts at integration time, the system can protect any new API, database, or web scraper the agent uses.
Sources
- Gemma 4 Technical Report
- Securing AI Agents with Information-Flow Control
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- Defeating Prompt Injections by Design
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
- Defending Against Indirect Prompt Injection Attacks With Spotlighting
- Preventing Prompt Injection with Type-Directed Privilege Separation
- Stronger Enforcement of Instruction Hierarchy via Augmented Intermediate Representations
- Prompt Flow Integrity to Prevent Privilege Escalation in LLM Agents
- DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM Agents
- AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
- Untrusted Content Masking for Web Agents with Security Guarantees
- PromptArmor: Simple yet Effective Prompt Injection Defenses
- Defending Against Prompt Injection with DataFilter
- Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
- Prompt Injection as Role Confusion
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs