PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents

arXiv:2610.01349 · cs.CR, cs.AI · Submitted 2026-10-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.

Elias: Today's paper: "PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents".

Nadia: Detailed Research Summary:

Elias: First, who's behind it and why it matters.

Paper summary: Nadia: To wrap up our discussion on "PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents," we see that the core contribution is shifting enforcement to the last tool boundary where every piece of evidence is available to commit to a verified contract.

Elias: I agree with Nadia; they’re focused on making that condition precise, which leaves the final decision point where an agent can still function while maintaining that security posture.

Priya: From my perspective, the impact seems to be less about a sudden leap in capability and more about establishing a formal structure for how we trust the dynamic interactions between an AI agent and its tools.

Nadia: Precisely, it’s about creating a verifiable layer of mediation that prevents the poisoning of tool metadata or skills from steering agent behavior later on.

Elias: The authors are showing how to separate the certified contract, PACE-C, from the evaluated configuration, PACE-P to handle restoration and repair without needing a full recertification every time.

Priya: That separation is interesting because it suggests that we don't need an impossibly strict gate for every single dynamic adjustment if we have a mechanism to handle repairs formally.

Nadia: Exactly, it’s about acknowledging the dynamic nature of these agents while still enforcing strict rules on what they are authorized to do at any given moment.

Elias: Ultimately, the paper is laying out a formal way to reason about noninterference and self-composition in this context so we can assess how far short of perfect security we actually are right now.

Conclusion: Nadia: So, we're wrapping up our discussion on "PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents," focusing now on what this title and its authors actually mean for us in plain terms.

Elias: I think the core concept boils down to giving an AI agent a really strict, verifiable way to check what tools it’s allowed to use at any given moment, which is a big deal for security.

Priya: From my side, I'm curious if this means we can actually start trusting these complex AI systems more because there's a formal structure behind how they execute their plans.

Nadia: Exactly, Priya; it’s about moving past just hoping the agent behaves correctly and having a mathematical framework to enforce those capabilities during runtime.

Elias: The authors are essentially proposing that the enforcement logic needs to be tied directly into the execution path itself so you can't easily tamper with what happens next.

Priya: That structural integrity is key for privacy researchers, because if we can prove how data flows through these tool calls, it gives us a much better picture of potential leakage points.

Nadia: I think the real implication here is that we can start designing AI agents with built-in security checks from the ground up instead of bolting on fixes later.

Elias: And that formal proof structure they mention suggests this isn't just a heuristic; it’s an attempt to build a solid foundation for how these agents manage their own actions securely.

Priya: It really looks like we might see better accountability as AI becomes more integrated into critical workflows, since there’s now a mechanism to track the provenance of those tool-generated actions.

Nadia: So, it boils down to making sure that when an agent proposes an action, its entire history and authority are checked before it actually gets executed.

Elias: That’s right; the paper lays out a method for creating a certified contract that governs every single tool call through this rigorous four-phase process.

Priya: It seems like this work could significantly impact how we audit AI systems, moving from reactive checks to proactive, verifiable security enforcement throughout the entire lifecycle of an agent's task.

Fengpeng Li, Qizhou Wang, Yuke Hu, Kemou Li, Jun Liu, Haiwei Wu

PRADA Lab, King Abdullah University of Science and Technology · Imperfect Information Learning Team, RIKEN Center for Advanced Intelligence Project · State Key Laboratory of Internet of Things for Smart City, University of Macau · National Institute of Informatics · School of Computer Science and Engineering, University of Electronic Science and Technology of China

cs.CR, cs.AI

Submitted: 2026-10-01

Updated: 2026-10-01

Code: https://github.com/meta-llama/PurpleLlama

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 83/100

The gist: This document synthesizes information from two distinct perspectives—a high-level technical overview of the PACE mechanism and a detailed artifact/experiment report—to provide a comprehensive

Key concepts

Provenance-Aware Capability Enforcement (PACE)
A security mechanism designed for LLM agents that enforces tool use by tracking the origin (provenance) of every action. It prevents agents from poisoning system components by rigorously checking proposed effects against a certified contract before execution.
Path Confinement (P)
This process involves creating an executable 'cut' in the agent's influence graph. It systematically severs any potentially malicious or unauthorized routes that could allow harmful actions to propagate through the agent's execution flow.
Certified Contract (PACE-C)
The PACE-C is a strict set of allowed operations verified during Phase II. It represents the guaranteed, safe contract for a specific tool call, ensuring that only authorized effects are permitted before the final decision is made.
Evaluated Configuration (PACE-P)
This configuration handles dynamic agent behavior like restoration and repair. While PACE-C sets the strict rules, PACE-P allows the system to restore blocked calls or dispatch repairs without needing a full re-verification.

Terminology

Summary

This document synthesizes information from two distinct perspectives—a high-level technical overview of the PACE mechanism and a detailed artifact/experiment report—to provide a comprehensive understanding of Provenance-Aware Capability Enforcement (PACE), a novel security framework designed for tool-using Large Language Model (LLM) agents.

The fundamental challenge addressed by PACE is the inherent risk posed by LLM agents that turn generated text into real side effects. This process allows for the poisoning of critical components—such as tool metadata, retrieved pages, memory states, and reusable skills—which can subsequently steer the agent's next actions. Traditional security measures like vetting an artifact before admission are insufficient because they cannot account for the dynamic nature of these side effects. Furthermore, a safe variant and a leaking variant can produce identical admission evidence; therefore, a sound gate cannot relax mediation at the egress site for either case.

PACE moves the critical decision point—the enforcement of tool calls—to the last enforceable tool boundary, where concrete proposals, authenticated authority, runtime provenance, and a working hook are all available simultaneously. It achieves this by mediating every tool call through a rigorous four-phase process:

  1. PROPOSE (Phase I): The agent proposes an action. During this phase, the system freezes the current arguments, extends the influence graph to incorporate schema-defined effects, and expands these potential effects based on the proposal.

  2. CUT & CERTIFY (Phase II): This is where path confinement and capability verification occur concurrently.

  • Path Confinement (P): Proposes an executable cut of represented influence paths, severing potentially malicious or unauthorized influence routes in the agent's execution graph.

  • Capability and Effect Verification (C): Checks the proposed effects against an authority compiled from the authenticated request. This distinguishes the certified contract (PACE-C)—the strict set of allowed operations—from the evaluated configuration (PACE-P), which handles restoration and repair behavior.

  1. ENFORCE (Phase III): The system installs the necessary manifest based on Phase II decisions and dispatches the call if it passes verification against PACE-C.

  2. FINALIZE (Phase IV): The final decision rule, defined by Equation (7) (dec(at) =), commits the outcome for the subsequent call, establishing a certified contract (PACE-C) that preserves the cut and ensures the final call passes effect verification.

PACE is grounded in formal proofs that establish strong security properties:

  • Represented-Path Separation (Proposition 1): If a path verifier accepts a certificate (cert t), removing the proposed cut (C t) ensures that the utility (U t) is unreachable from the current state (B t P t). This proves that path confinement effectively severs malicious influence.

  • Authority over Schema-Exposed Effect Atoms: The system guarantees authority over all effects exposed in the schema, linking concrete cuts to real execution links via executor grounding.

  • Transparent Chains Simulation (Proposition 2): This proposition ensures that if a sequence of calls adheres to specific host emission rules (reading/writing only at declared positions, preserving pass-through identifiers), the resulting state (gamma) can be successfully emitted and simulated in the system graph (G - t).

The framework clearly separates two operational modes:

  • PACE-C (Certified Contract): This is the contract that preserves path confinement (P) and final effect authorization. It represents the strict, verified set of allowed operations for a given call.

  • PACE-P (Evaluated Configuration): This configuration handles the dynamic aspects of agent behavior, specifically restoration and repair. The combiner associated with PACE-P can restore calls that were blocked by PACE-C and dispatch declared repairs without requiring a full final-call recertification.

The research rigorously validates PACE across eight executable agent-security benchmarks utilizing three target model families.

  • Performance: The full combination of P, C, and the final decision rule achieves the strictly lowest attack success on most benchmarks, achieving scores such as PACE 0.42 29.57 10.5 64.64 8.75 81.63 (Table 2).

Improvements for AI systems

Here are the specific improvements that can be made to AI systems by implementing the PACE framework, as described in this paper:

  1. Improve security against tool-using LLM agents by introducing a multi-phase enforcement mechanism (PACE) that mediates every tool call immediately before execution.

  2. Implement Path Confinement to propose executable cuts of represented influence paths within the agent's execution graph, ensuring that unsafe influence cannot reach terminal effects without explicit verification.

  3. Integrate Capability and Effect Verification to check schema-defined effects against an authority compiled from the authenticated request, ensuring that only authorized actions are permitted based on the request's capabilities.

  4. Distinguish between a Certified Execution Contract (PACE-C) and a Restoration/Repair Configuration (PACE-P), allowing the system to selectively restore blocked calls or apply declared repairs without compromising fundamental security guarantees.

  5. Establish a formal guarantee of Represented-Path Separation, ensuring that an attacker cannot reach an unsafe terminal effect if the required influence path has been severed by the enforcement mechanism.

  6. Ensure Authorized-Effect Soundness by verifying that every effect atom is covered by a valid capability and satisfies nine specific obligations (e.g., authorization coverage, confidentiality, persistence) before execution.

  7. Create a Conditional Grounded Execution Theorem, which guarantees that under specific conditions (current state, current epoch), any causal chain from a protected source to an unsafe effect cannot complete at the call site unless it traverses a path that is explicitly permitted by the enforcement mechanism.

  8. Enhance resilience against Trust Inversion and Admission-Time Evidence Confusion by precisely defining when admission-time evidence can or cannot license relaxing an egress site, preventing attackers from using self-written specifications to bypass security checks.

  9. Implement a formal, verifiable protocol (MAC chain) for runtime enforcement that ensures freshness, anti-replay capabilities (using one-shot authorization tokens), and durable execution state management across distributed deployments.

By implementing these improvements, the enhanced AI system can:

  1. Maintain strong isolation between model reasoning and real-world effects by strictly controlling what the LLM can do with its tools.

  2. Prevent prompt injection from steering subsequent tool calls or retrieving sensitive data through poisoned metadata or memory, as the influence paths are actively cut before execution.

  3. Guarantee that every action taken by a tool is authorized according to the authenticated request, mitigating unauthorized actions stemming from model-generated text manipulation of schemas or arguments.

  4. Provide a verifiable audit trail showing exactly which proposed actions were accepted and how they were modified (restored or repaired), allowing for high-assurance debugging and recovery.

  5. Enforce strict causality, ensuring that even if an attacker manages to construct a long sequence of steps, the enforcement mechanism will break the chain at the first point where an unsafe influence is attempted.

Abstract

Tool-using large language model (LLM) agents turn generated text into real side effects, so poisoned tool metadata, retrieved pages, memory, and reusable skills can steer the next call. Vetting an artifact before admission does not settle this. A safe variant and a leaking variant can produce the same admission evidence, and a sound gate then cannot relax that site for either. We make that condition precise, which leaves the last boundary a deployment can still act on. We present Provenance-Aware Capability Enforcement (PACE), which mediates every tool call immediately before it executes. Path confinement proposes an executable cut of represented influence paths, while capability and effect verification checks schema-defined effects against authority compiled from the authenticated request. We distinguish the certified execution contract from the evaluated configuration, which can restore an authorized call after a proposed block or apply a declared repair. Confinement requires the final action to preserve the certified cut. On eight executable agent-security benchmarks with three target-model families, the evaluated configuration gives strictly lowest attack success in 62 of 79 eligible attack columns and ties in 14; full-benchmark native utility loses at most three points relative to the undefended agent. A complete ablation over 1167 paired cases attributes most security gains to effect verification and refusal control to boundary adaptation. A reduced-scale adaptive search succeeds on 0/30 out-of-authority targets against the defense.

Sources

Related papers