ContractWarden: Kernel-Enforced Damage Boundaries for AI Agents via Human-Authorized Contracts

arXiv:2609.38248 · cs.CR, cs.OS · Submitted 2026-09-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.

Elias: Today's paper: "ContractWarden: Kernel-Enforced Damage Boundaries for AI Agents via Human-Authorized Contracts".

Nadia: Large language model agents can execute commands, create subprocesses, and directly access files and networks, allowing prompt injection or planning errors to become operating-system side effects.

Elias: First, who's behind it and why it matters.

Title and authors: Nadia: So we're diving into "ContractWarden: Kernel-Enforced Damage Boundaries for AI Agents via Human-Authorized Contracts," and it sounds like this paper is addressing the real danger of LLM agents messing with the operating system when they make mistakes or get tricked. It basically proposes a way to stop that before anything bad actually happens.

Elias: Exactly, Nadia, and what's really interesting is how they try to separate the idea of proposing a policy from actually enforcing it; this isn't just another application-level check where the agent could probably bypass it. They focus on binding a human-approved contract directly into the kernel's security mechanisms.

Priya: From a privacy standpoint, I'm curious about what kind of damage they are trying to prevent here—is it more about data exfiltration or something that messes with system integrity? Because if an agent can access files, we have to think about how that affects the data flow and what information could potentially be leaked.

Nadia: Well, the core idea is that instead of trusting the agent's suggestion on what's allowed, a human authorizer sets a tri-state contract—allow, deny, or no egress—and ContractWarden enforces it using an extended Berkeley Packet Filter Linux Security Modules data plane. That means the agent can only try to do what's explicitly in that contract.

Elias: That eBPF LSM enforcement is key because it moves the decision point deep into the kernel, which gives them a level of certainty that application-level checks just don't offer when you have a complex LLM planning process running. They’re using this data plane to "monotonically propagate no egress through processes, regular files, pipes, FIFOs, and supported Unix-domain sockets."

Priya: Monotonic propagation sounds robust for tracking state changes across different parts of the system; does that mean if a file gets marked as restricted on one end, that restriction sticks everywhere it goes? I want to know if there are any blind spots where a piece of data could sneak through without being tracked.

Nadia: They address that by defining the state of an object x as L(x), E(x), where L is a source bitmap and E is an egress-denial bit, and they define how these states join when processes derive or access objects. This formal state joining means the restrictions accumulate in a predictable way across all interactions.

Elias: I'm looking at the technical design here; they resolve assets to a "(device, inode) key" that contains "deny and no egress source bitmaps," and they specify how the kernel performs state updates like S(Q) from S(Q) S(P) on process derivation. That's where the cryptographic intuition meets the systems engineering of how to manage distributed state efficiently.

Title and authors: Priya: So, if they are using inode state for regular files and pipes, that gives them a clear picture of how data flows between different parts of the system; but I wonder what happens when things get complex, like with many interconnected services or symbolic links.

Nadia: They specifically mention that regular files, FIFOs, and anonymous pipes use inode state because Unix-stream peers have different inodes, so the sender joins state into a peer-state container and the receiver absorbs it. This is designed to handle those stream interactions where things can get messy.

Elias: And they put limits on policy keys too; they reject a contract that expands past four thousand ninety-six policy keys for directory compilation, and paths beyond thirty-two ancestors are outside the scope of their control plane metadata. These constraints help bound the complexity of what can be managed by this system.

Priya: Bounding the ancestor walk depth sounds like a practical measure against some kind of recursive or overly deep path traversal attempt, which is something I've seen in other security analyses where agents try to find hidden paths to sensitive locations.

Nadia: Exactly; and they also state that multiple tasks can safely share the long-lived data plane, and installation for one task preserves every other active source bit on the same object. This allows for efficient resource sharing while maintaining isolation at the enforcement layer.

Elias: That mechanism ensures that once a restriction is set on an object, it persists across different tasks using that shared data plane; it’s not just a temporary check for one specific execution path, which is much stronger than what we see in some simpler monitoring systems.

Priya: If the system can reliably track this state propagation, how does this translate into real-world protection against, say, an agent being tricked into opening a network connection it shouldn't? Is that covered by the egress denial bit?

Nadia: Yes, that's where the E(x) bit comes in; a no egress match allows access while setting E(P) = one in the same decision path, effectively denying any prohibited network effect before it occurs. This is what stops those direct side effects we see with prompt injection.

Elias: It's fascinating that they design the control plane to resolve assets into a "(device, inode) key" containing source bitmaps for deny and no egress, which then dictates the kernel action. This separation of proposal from enforcement seems like a very clean way to build trust without actually trusting the agent's internal logic.

Title and authors: Priya: I’m thinking about the implications for broader AI deployment; if we can guarantee that an agent operating within this framework cannot cause unauthorized file system changes, that opens up much more confidence for deploying agents in sensitive environments like research databases or even certain infrastructure management tasks.

Nadia: It really does, because they provide deterministic kernel enforcement for declared assets, which means the system operates exactly as the human authorized it to operate at the kernel level, rather than relying on the agent to behave correctly.

Elias: And when you look at their evaluation results, they show "deterministic kernel enforcement for declared assets" and passed all five hundred seventy security observations across tested properties. That level of coverage is substantial for a system dealing with LLM-driven actions.

Priya: It’s the kind of detailed testing that gives us confidence in the system's reliability, but I wonder if those tests fully capture every conceivable way an agent could try to probe or confuse the state joining mechanisms they described.

Nadia: The evaluation shows results across nineteen security tests satisfying predefined return-value and side-effect criteria, demonstrating controlled object lifecycles. They specifically tested things like "Source-isolated deny" and "Post-infection established send Send succeeds before and fails after infection," which speaks to the propagation resistance we discussed.

Elias: That level of rigor in testing the propagation—ensuring that state laundering is genuinely resisted by their design—is what separates this approach from systems where a single path might lead to a state change being ignored later on. It’s about ensuring the no egress propagates reliably through pipes and files.

Priya: So, looking at the overall picture of ContractWarden, it seems the paper is focused on creating a verifiable security boundary around agent actions by shifting trust away from the agent's policy generation and placing it firmly in human-defined kernel constraints.

Nadia: That’s right; they turn a human-authorized asset contract into concrete kernel state through gate-before-exec registration, which is the mechanism that locks down execution before any untrusted code runs. It’s about making sure the setup itself is solid before the agent gets a chance to cause trouble.

Elias: And it's that pre-execution binding that prevents things like "Pre-sideeffect denial returns EPERM before a prohibited asset or network effect," which means we catch the error at the lowest possible level, long before any actual damage can occur. That’s a crucial point for cryptographic integrity in this context.

Priya: I think the most significant implication for me is how this could apply to federated learning scenarios where agents might be used to manage data access; if we can ensure an agent cannot unilaterally violate a privacy contract enforced at the kernel level, that really elevates the security of those distributed training processes.

Title and authors: Nadia: That’s a strong point, Priya; because it moves the guarantee from "the agent followed its rules" to "the kernel enforces the human's rules," which is a fundamentally different and much more reliable security posture.

Elias: And from a cryptographic viewpoint, they are using this structure to model processes and kernel objects as graph nodes and interactions as directed edges, which aligns nicely with how we think about verifiable information flow models across complex computational graphs.

Priya: It sounds like the work moves security from a reactive auditing phase to a proactive, preventative enforcement mechanism embedded directly into the operating system's core security primitives.

Nadia: Precisely; they are providing a reference monitor that doesn't just observe and report; it actively denies prohibited side effects based on those human-defined contracts. This is what makes ContractWarden different from simply running an agent in a sandbox.

Elias: The performance metrics they provide, showing median overhead of "eleven point nine six–twelve point eight nine percent in a Linux six point one five virtual machine and thirty-five point seven nine–sixty-one point five four percent on a Linux six point one five physical platform," suggest they are achieving deterministic kernel enforcement even when running on bare metal, which is a good sign for practical deployment feasibility.

Priya: That performance gap between virtual and physical platforms is telling; it suggests the overhead of this eBPF mechanism is manageable enough to be considered viable for real-world systems, even if the physical platform runs a bit slower than the VM in these specific tests.

Nadia: It shows that deterministic kernel enforcement isn't just theoretical; it’s being measured under realistic workloads, and they managed to keep the overhead within a range where it's considered controlled. This moves it from lab curiosity toward something you might actually deploy in production environments where agents are running.

Elias: So, to wrap up on this paper, "ContractWarden: Kernel-Enforced Damage Boundaries for AI Agents via Human-Authorized Contracts," we see a system that uses eBPF to bind human contracts to kernel state via gate-before-exec registration and monotonic propagation of no egress decisions.

Priya: I think the main implication is establishing a much stronger, verifiable security guarantee for agents by ensuring their actions are strictly confined to pre-vetted operational boundaries defined by humans.

Nadia: That's right; it provides a way to reliably execute complex tasks within strictly defined, human-vetted security boundaries, ensuring that even if the LLM agent is compromised or makes a planning mistake, its actions are confined to the specific assets and network paths explicitly permitted by a trusted human contract.

The paper's summary: Nadia: So, to recap, ContractWarden is essentially a way to put a human's security contract directly into the Linux kernel using eBPF so that an AI agent can't just wander off and cause trouble without explicit permission.

Elias: That’s right, Nadia, and it really boils down to binding a high-level policy idea—the allow, deny, or no egress tri-state contract—to concrete kernel state before any untrusted code actually gets a chance to run.

Priya: From the perspective of measurement and privacy research, what I find compelling is how they formalize that "no egress" concept; it seems like a very specific way to track data flow denial across pipes and files in a way that’s hard to bypass.

Nadia: Exactly, Priya, because they use this monotonic propagation through process derivation and object access, which means once an egress restriction is set on a file or process, that restriction sticks everywhere it goes without getting lost in the system's complexity.

Elias: That propagation mechanism is what makes them resistant to state laundering; it ensures that even if an agent tries to move data through several intermediate pipes or sockets, the denial bit gets tracked along the way.

Priya: I’m interested in how this relates to real-world data privacy; if an AI agent is meant to process sensitive information, this kernel enforcement gives us a verifiable promise that the data won't leave its designated area without a human-approved reason.

Nadia: It does, Priya, and they’ve shown through those rigorous tests that this approach yields deterministic kernel enforcement for the assets that are declared, which means we know exactly what the system is allowed to touch at the lowest level.

Elias: And when you look at how they handle pre-execution binding—making sure validation and installation actually finish before execution begins—that’s a crucial safeguard against side effects appearing unexpectedly during setup.

Priya: That pre-execution check seems vital for any system dealing with complex workflows, because it stops the risk of an agent executing something dangerous just because its planning process got stuck mid-way through configuration.

Nadia: It’s about ensuring that the setup itself is solid before we let the agent do anything else, which is a lot better than hoping an application-level check catches everything in time.

Elias: And looking at the performance results, even on physical platforms, they show overhead that’s remarkably low compared to some of those ActPlane baselines they tested earlier, which suggests this isn't just a theoretical concept but something practical for deployment.

Priya: That deterministic behavior across different hardware setups is what really gives me confidence in the data shown in their five hundred seventy security observations, indicating that the system behaves predictably under stress.

Nadia: It’s about moving from trusting the AI's internal logic to trusting a formally defined and verified kernel boundary, which is a significant step forward for agent security.

Elias: And thinking about the future work they mentioned, especially generation-safe source reuse and atomic policy publication, that points toward making this mechanism even more scalable across different types of agents.

Priya: That scaling potential is what really excites me; if this framework can handle more complex concurrent agent workloads while maintaining this level of strict state tracking, it opens up new possibilities for secure AI-driven infrastructure management.

Nadia: Indeed, Priya, and I think the real world impact here is providing a foundation where we can reliably deploy AI agents in sensitive environments because we have a verifiable kernel guarantee that their actions stay within the human-defined limits.

The paper's improvements: Nadia: So, to summarize the paper's suggested improvements, ContractWarden isn't just about what it does now; it’s about making it even tougher by focusing on how we handle source reuse and policy updates.

Elias: That's right, Nadia, they are looking at ways to make the policy publication atomic so that there are no gaps or partial updates in the kernel state when a contract is changed.

Priya: From a privacy angle, I’m curious about how generation-safe source reuse addresses concerns about data leakage across different sessions or tasks; if an agent reuses a source identifier, does that risk exposing previously restricted information?

Nadia: It directly tackles that with generation-safe source reuse, meaning the system needs to be able to track which sources are reused safely without violating the integrity of the established damage boundary.

Elias: And atomic policy publication is important for cryptographic security because it ensures that a state transition—like changing an egress denial—happens completely or not at all, preventing intermediate states that could be exploited by a sophisticated attacker.

Priya: The authors also touch on trusted declassification, which seems like a big deal for AI systems that might process data from multiple sources; it implies a more granular control over what information is actually accessible based on the human contract.

Nadia: Exactly, Priya, and this level of control is what moves us closer to making these agents reliable tools in sensitive domains because we're not just relying on the agent’s current state but a formally defined, persistent security model.

Elias: I see how generation-safe reuse combined with atomic updates helps reinforce the monotonic propagation they established earlier, ensuring that the kernel's enforcement remains consistent even when the policy itself is being actively modified.

Priya: If these improvements hold up under stress testing for concurrent agent workloads, it could mean we can deploy these systems in environments where agents are constantly changing their operational parameters without compromising security.

Nadia: That's the goal, Priya; making sure that the core enforcement mechanism stays robust even when the AI agent is dynamically evolving its behavior.

Elias: The implication for me is that they are moving toward a system where the cryptographic assumptions around state joining are much more tightly bound to real-world operational constraints rather than just theoretical models.

Priya: It seems like these refinements give us a better picture of how to achieve system-level privacy guarantees, rather than just model-level protections for individual components.

Nadia: So, the core idea is building a highly resilient security layer that evolves alongside the agent’s tasks while remaining strictly tethered to its initial human authorization.

Conclusion: Nadia: So, to wrap up our discussion on ContractWarden: Kernel-Enforced Damage Boundaries for AI Agents via Human-Authorized Contracts, the main takeaway is that we’ve developed a mechanism to enforce human-defined security contracts directly at the kernel level.

Elias: Exactly; it’s about turning an abstract policy idea into concrete, verifiable kernel state through gate-before-exec binding and monotonic propagation of those no egress decisions.

Priya: I think the real impact here is establishing a much stronger, verifiable security posture for AI agents by ensuring their actions are strictly confined to pre-vetted operational boundaries defined by humans.

Nadia: And it does, Priya; this provides a way to reliably execute complex tasks within those human-vetted security limits, even when the underlying AI agent is making dynamic decisions.

Elias: Looking at the results, they show deterministic kernel enforcement for declared assets and passed all five hundred seventy security observations across their tested properties, which speaks to a high degree of reliability in that enforcement.

Priya: That level of coverage across those nineteen security tests is substantial; it suggests the system’s behavior under stress is very well understood, which is vital for privacy researchers trying to model the actual data flow risks.

Nadia: It really does; and I think the most significant implication is that we're shifting trust away from the AI's internal logic and placing it firmly in a verifiable kernel constraint managed by a human authorizer.

Elias: That’s right, Nadia, and when you consider how they handle state joining across process derivations and object access, you see how the cryptographic assumptions align perfectly with the systems engineering of managing distributed state efficiently.

Priya: I just think that moving security from an auditing phase to a proactive, preventative enforcement mechanism embedded in the operating system core is a big step toward securing complex AI deployments in sensitive areas.

Nadia: It’s about providing a reference monitor that doesn't just observe and report; it actively denies prohibited side effects based on those human-defined contracts, which is what makes ContractWarden distinct from simple sandboxing.

Elias: And the performance metrics they shared, showing controlled overhead even on physical platforms, tell us this isn't just a lab curiosity but something that could be practically deployed in environments where agents are running for extended periods.

Priya: Those practical deployment numbers give me confidence in the data; it suggests that the security guarantees aren't only theoretical but also feasible when dealing with real-world workloads.

Nadia: It’s a solid piece of work, and I think we need to keep an eye on those future plans they mentioned regarding generation-safe source reuse and atomic policy publication.

Elias: Those future steps are where the real scaling potential lies, as they aim to make the system even more resilient against evolving agent behaviors over time.

Priya: So, ContractWarden is really showing us how to build a foundation where AI agents can operate securely within strict privacy contracts defined by human oversight.

Dongxu Cui, Zhichao Gu, Ping Zheng, Wenshuai Xi, Simeng Han, Yong Liao

School of Cyber Science and Technology, University of Science and Technology of China · China Greatwall Technology Group Co., Ltd.

cs.CR, cs.OS

Submitted: 2026-09-29

Updated: 2026-09-29

Comments: Supersedes the preprint DOI 10.21203/rs.3.rs-10865359/v1, which has a different title. Submitted to ICOIN 2027. 6 pages, 2 figures

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 78/100

The gist: Large language model agents can execute commands, create subprocesses, and directly access files and networks, allowing prompt injection or planning errors to become operating-system side effects.

Key concepts

Tri-state Asset Contract
A contract proposed by a model that dictates whether an asset should be allowed, denied, or if egress is prohibited. Crucially, this proposal is not final; a human must make the ultimate decision. This step separates the AI's suggestion from the actual security policy.
Extended Berkeley Packet Filter (eBPF) LSM Data Plane
This is a Linux security module mechanism used to enforce decisions at the kernel level for file and network operations. It allows ContractWarden to monitor and control every interaction with files and sockets, ensuring that the contract's rules are strictly followed during execution.
State Join Monotonicity
A rule defining how asset states combine when a process or object is derived from others. The state joins based on combining source bitmaps (L) and egress-denial bits (E), ensuring that access decisions accumulate correctly and predictably across the system.

Terminology

Summary

Large language model agents can execute commands, create subprocesses, and directly access files and networks, allowing prompt injection or planning errors to become operating-system side effects. ContractWarden presents a Linux reference monitor that enforces a human-authorized damage boundary without trusting the agent or its policy suggestions.

How it works

ContractWarden separates policy proposal, semantic authorization, and privileged installation. A model only proposes a tri-state asset contract—allow, deny, or no egress—but a human makes the final choice. An execution gate binds this contract to a concrete task before untrusted code runs. The system relies on an extended Berkeley Packet Filter (eBPF) Linux Security Modules (LSM) data plane to enforce file and network decisions and monotonically propagate no egress through processes, regular files, pipes, FIFOs, and supported Unix-domain sockets.

System Model and Design

The system involves a human authorizer H who maintains a declared asset set R. For any task τ, the confirmed contract is the total function Γτ: R → allow, deny, or no egress. The state of an object x is defined as S(x) = ⟨L(x), E(x)⟩, where L is a source bitmap and E is an egress-denial bit. State joins monotonically: ⟨L1, E1⟩ ⊔ ⟨L2, E2⟩ = ⟨L1 ∨ L2, E1 ∨ E2⟩. The design targets six security invariants: Closed authorization requires every enabled asset to appear exactly once and prevents a model suggestion from directly becoming policy, and Pre-execution binding prevents untrusted code from running before installation and registration complete.

Authorization and Pre-Execution Binding

The control plane validates actions against the human's declared assets, rejecting missing, duplicate, unknown, or out-of-domain actions. The launcher creates a child blocked on a dedicated file descriptor to obtain a pidfd. Only after the daemon installs the asset policy and attaches task state does it open the execution gate. This mechanism ensures that Pre-sideeffect denial returns EPERM before a prohibited asset or network effect, and fail-closed setup keeps the child behind its gate if validation, installation, registration, or read-back fails.

Compilation, Propagation, and Denial

The control plane resolves assets to a (device, inode) key, where the value contains deny and no egress source bitmaps. State joins across process derivation and object access: S(Q) ← S(Q) ⊔ S(P) on process derivation, S(O) ← S(O) ⊔ S(P) on writable access, and S(P) ← S(P) ⊔ S(O) on readable access. A supported file operation is denied when the task’s sources intersect the target inode’s deny bitmap. A no egress match allows access while setting E(P) = 1 in the same decision path.

Performance and Evaluation

The system was evaluated across 19 security tests, satisfying predefined return-value and side-effect criteria. The results demonstrate deterministic kernel enforcement for declared assets, supported paths, and controlled object lifecycles. Median overhead on three co-located file-I/O workloads was 11.96–12.89% in a Linux 6.15 virtual machine and 35.79–61.54% on a Linux 6.15 physical platform, which is lower than the evaluated frozen ActPlane baseline for those specific runs, showing deterministic kernel enforcement. The system passed all 570 security observations across the tested properties, including tests for Source-isolated deny, Post-infection established send Send succeeds before and fails after infection, and Directory ancestor bound Depth 31 matches; depth 32 is outside scope.

Conclusion

ContractWarden turns a human-authorized asset contract into task-bound kernel state. It achieves this through gate-before-exec registration, object-level monotonic propagation to resist state laundering, and file plus multi-stage network hooks to deny prohibited side effects. The work concludes that All 570 security observations pass, while showing lower overhead than the frozen ActPlane configuration in co-located tests. Future work will focus on adding generation-safe source reuse, atomic policy publication, trusted declassification, and realistic concurrent agent workloads.

The gist: ContractWarden is a Linux reference monitor that enforces a human-authorized damage boundary without trusting the agent or its policy suggestions. It achieves this by binding a human-approved asset contract to a concrete kernel task using an eBPF LSM data plane to enforce file and network decisions monotonically.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems based on ContractWarden, and what those improved systems can achieve:

  1. The AI system gains a robust, human-authorized damage boundary enforced directly by the Linux kernel via an eBPF LSM data plane, rather than relying solely on application-layer validation or agent policy suggestions.

  2. The system can be guaranteed to enforce a pre-defined asset contract (allow, deny, or no egress) before any untrusted code executes, eliminating the risk of prompt injection or planning errors leading to unintended operating-system side effects (like unauthorized file access or network connections).

  3. The AI agent can propose a policy via a tri-state contract over declared assets, but the human retains the final trusted authorization root, ensuring that kernel enforcement is always bound to a concrete, verified task.

  4. The system achieves state laundering resistance through gate-before-exec binding and object-level monotonic propagation, meaning that once an egress restriction is set on a file or process, it cannot be bypassed by intermediate files, pipes, or IPC channels.

  5. The AI agent's execution environment gains deterministic kernel enforcement for declared assets and controlled object lifecycles (e.g., bounded directory depth limits), preventing resource exhaustion attacks stemming from planning errors.

In summary, the improved AI system can reliably execute complex tasks within strictly defined, human-vetted security boundaries, ensuring that even if the LLM agent is compromised or makes a planning mistake, its actions are confined to the specific assets and network paths explicitly permitted by a trusted human contract.

Sources

Related papers