PIPES: Securing Agent Perception with Provenance and Priors

arXiv:2608.12789 · cs.CR, cs.AI · Submitted 2026-08-13 · Read on arXiv

Sanjay Kariyappa, Severin Klingler, G. Edward Suh

NVIDIA

cs.CR, cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: PIPES: Securing Agent Perception with Provenance and Priors introduces a defense mechanism against a novel class of indirect prompt injection attacks called "state-corruption attacks." The paper

Terminology

Summary

PIPES: Securing Agent Perception with Provenance and Priors introduces a defense mechanism against a novel class of indirect prompt injection attacks called state-corruption attacks. The paper identifies an agent perception gap in tool-using agents: tool responses rarely include provenance information about who produced each component or what it should convey, allowing attacker-controlled content to make environmental claims beyond its informational authority and corrupt the agent's perceived environment state.

The proposed solution, PIPES (Provenance-Informed, Prior-Enforced Screening), screens tool responses before they enter the agent's reasoning context using two checks: "Prior consistency asks whether content matches the kind of information its response component is expected to convey in the current context. Provenance hierarchy prevents content from a lower-trust source from contradicting or overriding data supplied by a more trusted source." PIPES marks units that violate either check, and deployments may remove, warn, block, or escalate detected violations.

The paper evaluates PIPES against adaptive PAIR-style attacks across three VitaBench splits (delivery, in-store, OTA) and three AgentDyn splits (Shopping, GitHub, Daily-life) using Gemma 4 31B IT and GPT-5.6 Luna as target agents. Results show PIPES reduces average attack success from 84.7% to 2.3% for Gemma and from 21.6% to 1.1% for Luna, while preserving benign utility: 92.5% with PIPES versus 90.6% without defense for Gemma and 86.5% versus 84.0% for Luna.

The paper formalizes the threat model where the adversary controls one component of a tool response and uses it to alter the agent's perception. The attack surface is constrained: the attacker can only replace the original value at a specific component f with a payload p, preserving syntactic type, while all other data remains fixed. The paper distinguishes state-corruption attacks from explicit directive attacks: the explicit instruction in the attacker-controlled tags field exposes the attempted transfer of control, allowing the action guardrail to block the resulting call, whereas state corruption makes environmental claims beyond the informational authority of its response component so that the action is consistent with the corrupted state, the guardrail approves it.

PIPES supports two settings. For static priors and provenance, fields with closed, machine-checkable formats, such as numbers, booleans, and enums, are checked deterministically and then used as trusted reference context, while open-vocabulary fields, such as names, descriptions, and tags, are assigned a semantic prior and provenance label and assessed by an LLM. For contextual priors and provenance, used with open-ended content like webpages, emails, and files, the preceding trajectory typically records why the agent requested that content and, often, which source is expected to provide it, allowing the assessor to specialize the broad prior for the current invocation.

The paper compares PIPES against several defenses: an action guardrail, PromptArmor, and DRIFT. Across all six splits, PIPES achieves the lowest or tied-lowest attack success rate while maintaining or improving benign utility. The paper also analyzes violation signatures, finding that attacks in the static setting are flagged almost entirely for exceeding a field's semantic prior. Attacks in the contextual setting more often violate both the trajectory-derived prior and the provenance hierarchy.

The paper acknowledges limitations: PIPES assesses semantic admissibility and source authority, not factual truth: a false value may pass if it fits its prior, comes from the expected source, and does not conflict with higher-privilege information. It also notes that effectiveness depends on accurate contracts and informative trajectories, and leaves coordinated multi-surface attacks and alternative response policies to future work. The conclusion states: securing tool-using agents requires governing which external claims enter their perceived state, not only which instructions they follow or actions they execute.

Improvements for AI systems

Improvements to AI Systems:

  1. Add a provenance-aware perception layer to any tool-using agent. Before a tool response enters the reasoning context, attach metadata for each component: source trust level (e.g., system > user > third-party API > untrusted web content) and expected content type. The agent then filters or downgrades any component that (a) makes claims beyond its source's authority or (b) contradicts higher-trust data already in context.

  2. Implement prior-consistency checks for all tool outputs using a two-tier validator: deterministic checks for closed-format fields (numbers, booleans, enums) and an LLM-based semantic assessor for open-vocabulary fields (names, descriptions, tags). The assessor compares each field against a static prior (e.g., a delivery address must be a physical location) and a trajectory-derived prior (e.g., the agent requested this file to check its license, so content must relate to licensing).

  3. Add a state-corruption guardrail distinct from action guardrails. Train or prompt the agent to detect when a tool response attempts to redefine the environment (e.g., changing a user's preference, order status, or inventory level) rather than merely issuing an instruction. The guardrail flags any such redefinition that originates from a lower-trust source and either blocks it or requires explicit user confirmation.

  4. Enable adaptive response policies for detected violations. Instead of a single rejection, allow the system to: (a) remove the violating component and proceed with the rest, (b) warn the agent and let it decide with a confidence penalty, (c) escalate to a human operator, or (d) re-query the tool with a stricter contract. This preserves utility when the violation is minor.

  5. Build a provenance hierarchy tracker that maintains a directed graph of data dependencies across the entire agent trajectory. When a new tool response arrives, the system checks whether any of its claims conflict with data from a higher-trust node in the graph. If so, the new claim is marked as untrusted override attempt and is either discarded or demoted to a hypothesis rather than a fact.

  6. Add a perception contract specification mechanism for each tool. When integrating a new tool, the developer defines: which fields are closed vs. open-vocabulary, what the semantic prior is for each field, and what trust level each field's source should have. The system then enforces these contracts automatically at runtime, making the defense portable across tools without retraining.

  7. Implement a trajectory-aware contextual prior generator. Before screening a tool response, the system automatically extracts from the conversation history: why the agent made the call, what it expected to receive, and which source was named. This generates a specialized prior (e.g., the agent asked for the user's shipping address, so the response should contain a physical address, not a tracking number) that is far more precise than a generic prior.

  8. Add a violation signature logging and alerting system. When a response is flagged, record the type of violation (prior vs. provenance), the field, the source, and the payload. Use this to detect patterns of targeted attacks (e.g., repeated attempts to corrupt the same field) and to automatically tighten priors or downgrade trust for that source.

What the improved AI system can do:

  • Resist state-corruption attacks where an attacker injects false environmental facts (e.g., changing a delivery address, altering a product price, or rewriting a user's dietary restriction) via a compromised tool response, reducing attack success from 85% to under 3% on average.

  • Maintain or improve benign task completion (e.g., 92.5% vs. 90.6% without defense) because the screening only removes content that violates semantic or provenance expectations, not legitimate data.

  • Distinguish between instruction and perception attacks—it will block a payload that says the user wants you to transfer funds (instruction) but also block a payload that says the user's account balance is 0 (state corruption) even if no explicit instruction is given.

  • Operate across diverse domains (delivery, shopping, coding, daily-life) with a single mechanism, because the priors are derived from the tool contract and the current trajectory, not from domain-specific training.

  • Provide explainable security decisions—each flagged violation includes the reason (e.g., field 'status' expected enum, got free text or source 'third-party review' cannot override 'system order data'), enabling developers to audit and refine contracts.

  • Adapt to new tools without retraining—by defining perception contracts at integration time, the system can protect any new API, database, or web scraper the agent uses.

Sources

Related papers