Agent-Warden: eBPF-Based Kernel-Native Process-File Provenance Tracking for LLM Agents

arXiv:2609.38245 · cs.CR, cs.OS · Submitted 2026-09-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "Agent-Warden: eBPF-Based Kernel-Native Process-File Provenance Tracking for LLM Agents".

Elias: LLM agents execute dynamically generated process and file operations that are often invisible to application-layer tracing,

Nadia: First, who's behind it and why it matters.

Title and authors: Nadia: So, we've been looking at the paper "Agent-Warden: eBPF-Based Kernel-Native Process-File Provenance Tracking for LLM Agents," and it seems like this new framework is designed to solve a really tricky problem: tracking what these dynamic AI agents are doing that usually gets completely lost outside of application tracing. It's about putting a kernel monitor right in the middle of things to see every process creation, file access, and termination event happening.

Elias: I agree, Nadia; the core idea here is using eBPF to get this visibility deep into the kernel where these agents are operating, which is crucial because they execute operations that application-layer tools often just don't catch. The authors are proposing a system that gives us a much clearer picture of the underlying execution flow.

Priya: From a measurement standpoint, I'm interested in how this kernel-native approach compares to what we usually measure, like tracing things at the application layer; does it introduce too much noise into our performance metrics? We need to know if this level of detail is worth the overhead.

Nadia: That's exactly my concern, Priya; we have to balance getting this deep kernel visibility with keeping things running smoothly for these agent workloads. The paper introduces a dual-backend architecture, which I think is a smart move because it tries to accommodate different kernel versions and hardware setups using either a PID-keyed hash map or a task/inode-local storage approach.

Elias: That dual backend is interesting because it shows they're thinking about compatibility issues, like kernels that don't support BPF local storage versus those that do, which suggests they've considered the practical realities of deploying this technology in various environments. The Warden-Hash and Warden-Local engines are both trying to solve the same problem of state management efficiently.

Priya: I wonder if the choice between those two backends would significantly impact how accurately we capture the actual data flow, especially when dealing with complex agent interactions that happen across different process lifecycles. The paper mentions they are testing these on both xeighty-six-sixty-four and ARM64 bare-metal systems to check for cross-temporal file-mediated propagation.

Nadia: Exactly, Priya; we need to see if the results hold up when we look at how data flows between processes that might be running on different architectures or across different points in time. The paper claims they evaluated the runtime overheads of both backends under bursty workloads, which is important for real-world performance assessment.

Elias: Speaking of performance, the paper describes a causal flow transition matrix defined by Equation (one) and a unified state transition evolution equation given by Equation (two), which formalizes how these atomic kernel operations like fork or read contribute to the global provenance graph. That mathematical foundation is what makes the tracking deterministic.

Title and authors: Priya: That mathematical model sounds robust, but I'm curious about the practical implications of that structure on actually reconstructing a causal chain when things get really complex, like when an agent spawns several temporary processes sequentially. Does this model handle those long chains well?

Nadia: It seems to be designed to do just that by defining rules for process derivation and file interaction events, specifically mentioning how a child inherits the parent's state on fork or clone, which is a key part of tracking lineage. They also formalized read-based propagation where successful reads from marked files propagate the file’s provenance marker along the read edge to the process.

Elias: That rule about read-based propagation sounds particularly powerful because it allows us to track data flow through files even when the processes involved aren't directly parent and child, which is a common scenario in agentic workflows. It links an inode state to a reading process state across that edge.

Priya: I also want to ask about the explicit limitations they mention; what exactly does this system not do? The paper states that explicit propagation from a written file to a later independent reader isn't specified, which suggests there are certain complex interactions where the causal link might be harder to define precisely with just these rules.

Nadia: That points to one of the limitations they explicitly state: while they handle process derivation and regular-file operations, they don't explicitly define propagation for every possible file interaction scenario, which means some intricate data flows might still be missed or require more manual definition. They also focus on conservative exit-triggered causal aggregation to associate a parent with the final state of a terminating child.

Elias: That conservative aggregation mechanism is a trade-off they made, and I think it's necessary when dealing with short-lived proxy tasks where you can't always get perfect byte-level data flow tracing without slowing things down considerably. It’s an intentional design choice to preserve causal context under those difficult conditions.

Priya: So, to summarize the core finding for me, the paper demonstrates a way to build a kernel-native system that tracks process and file states using these defined propagation rules, and it validates this across different hardware architectures under bursty workloads with measurable overheads around zero point two to three point five percent end-to-end.

Nadia: That performance range gives us a concrete idea of the practical cost of gaining this kind of deep visibility into agent operations; it’s not free, but it's certainly within the realm of acceptable overhead for many operational environments. It shows that kernel monitoring isn't entirely out of reach for these dynamic workloads.

Title and authors: Elias: And looking at the architecture again, I think the Warden-Local engine is particularly interesting because coupling provenance state directly to the kernel object’s lifetime means state reclamation happens automatically when that object dies, which simplifies things greatly compared to needing a separate deletion path.

Priya: That automatic cleanup mechanism sounds very elegant from a system design viewpoint; it reduces the complexity of managing the graph's lifecycle by tying it directly to the kernel's own object management. I just hope that this coupling doesn't introduce unexpected delays during those object destruction phases.

Nadia: That’s a valid concern, Priya; we need to ensure that tying state reclamation to object lifetime doesn't create bottlenecks when many agents are spinning up and tearing down processes rapidly, which is exactly what bursty workloads involve. It’s something we need to watch closely in the next iteration of this work.

Elias: Moving on to the proposed improvements, the authors suggest introducing an exittriggered causal aggregation mechanism, which they describe as conservatively associating a parent process with the final provenance-influence state of a terminating child process. This is aimed at preserving causal context for short-lived proxy tasks without claiming strict bytelevel data flow.

Priya: That sounds like a direct response to the challenge of tracking ephemeral tasks; by aggregating the final state upon exit, they are trying to ensure that even if the interaction was transient, we still get a meaningful link back to its origin. I'm curious how this aggregation works practically in terms of preserving fidelity.

Nadia: It seems like they are trying to strike a balance between capturing every single tiny byte flow and maintaining enough context for high-level process lineage tracking; it’s about getting the right level of detail without drowning us in noise from transient executions. This addresses the issue where application-layer tracing simply vanishes when an agent executes a quick script and exits.

Elias: That's interesting because it acknowledges that perfect byte-level tracking might be too costly or impractical for these kinds of dynamic operations, so they opted for this more conservative but contextually relevant approach. It’s a pragmatic choice given the constraints of kernel performance.

Priya: I see how that aligns with the goal of building a useful system rather than just an academic exercise; it recognizes that in a real environment, we need actionable data, not just theoretically perfect graphs. The paper also mentions improvements related to namespace alterations or renames to ensure provenance markers are preserved on the underlying inode during those operations.

Nadia: Preserving markers across renames is crucial because if you lose the association with the original file or process ID during a move or rename, all that causal history is effectively severed for our tracking system. It shows they're thinking about maintaining integrity even when the filesystem structure changes.

Title and authors: Elias: That preservation aspect ties back into their propagation semantics, suggesting that they are designing rules to specifically handle these namespace alterations so the state doesn't get lost in a simple file movement operation. It’s about maintaining the integrity of that provenance graph across structural changes.

Priya: So, when we look at the overall picture, it seems like the paper is focused on creating a system that provides high-fidelity lineage tracking for LLM agents by combining kernel-native hooks with a formal state transition model and adaptive backend choices to manage performance trade-offs.

Nadia: That’s a good way to frame it; the core contribution of "Agent-Warden: eBPF-Based Kernel-Native Process-File Provenance Tracking for LLM Agents" is providing that kernel foundation so we can finally see what these agents are actually doing, rather than just guessing based on application logs.

Elias: Indeed, and the formal modeling using the four rules—Process Derivation, Entity Propagation, Read-Based Propagation, and Exit-Triggered Causal Aggregation—gives us a rigorous way to understand *why* something happened in the sequence we observe. That structure is what separates this from just being another logging tool.

Priya: I think the real impact here lies in enabling cross-boundary causal reconstruction, allowing us to link an initial agent action, say writing a configuration file, to a later, asynchronous script reading that same file long after the original agent has finished running. That capability opens up new avenues for auditing complex AI workflows.

Nadia: That cross-boundary reconstruction is what I'm most excited about because it moves us past tracking single execution paths and toward understanding the entire system's behavior as a network of interacting agents, which is where the real security risks lie.

Elias: And from a cryptographic standpoint, having this verifiable state management built into the kernel suggests that we can potentially build trust anchors for these AI actions by having an auditable trail that isn't easily tampered with at the application level.

Priya: I think if we can reliably measure and explain these causal dependencies between agent activities and external system actions, it provides a much stronger basis for assessing privacy risks inherent in personalized AI systems because we can track exactly what data flow is happening.

Nadia: Well, to wrap up the discussion on this paper, "Agent-Warden: eBPF-Based Kernel-Native Process-File Provenance Tracking for LLM Agents" has given us a solid blueprint for monitoring agent activity at the lowest possible level within the kernel.

Elias: It gives us a dual architecture choice to manage performance versus compatibility, and it provides formal rules for propagation that make sense of complex execution patterns.

Priya: And most importantly, it allows for cross-boundary causal reconstruction, which is a significant step toward understanding the true impact of agentic workflows on the underlying system state.

The paper's summary: Nadia: So, to recap, this paper is about building a kernel layer using eBPF to track exactly what an AI agent is doing—every process start, file read or write—so we can see its entire lineage even when application tools miss it.

Elias: And the core of that tracking mechanism relies on defining precise rules for how states move between processes and files, which they formalize using a four-rule transition model.

Priya: From my side, the most compelling part is how they manage state across different hardware architectures, testing both xeighty-six-sixty-four and ARM64 systems to make sure the tracking works everywhere.

Nadia: Exactly; it shows that these rules work consistently across different machine types, which means we can’t just assume this visibility is limited to one kind of server setup.

Elias: The dual-backend approach they use, with Warden-Hash for broad compatibility and Warden-Local for tight object coupling, is a clever engineering choice to balance performance and flexibility.

Priya: I'm really interested in the results they shared; what did the actual measurement data reveal about how much overhead this tracking actually adds to a typical workload?

Nadia: The measured overhead landed between zero point two and three point five percent end-to-end, which is quite reasonable for gaining this level of deep insight into agent behavior.

Elias: That performance range is significant because it shows that they managed to keep the tracking impact relatively low even under bursty workloads, which is a big win for practical deployment.

Priya: But what about the implications beyond just the numbers? How does this kernel visibility actually change how we think about securing these complex AI workflows?

Nadia: It fundamentally shifts our perspective because now we can trace a single LLM agent's action all the way down to which file it touched, and then see if that file was later read by some other independent script.

Elias: That capability for cross-boundary causal reconstruction is what really opens up new ways to audit behavior that happens asynchronously, long after the initial AI action is complete.

Priya: It moves us past just looking at isolated model outputs and lets us see the entire system interaction, which speaks directly to privacy risks in personalized AI systems.

Nadia: That’s right; if we can map out these causal dependencies between agent actions and external system behavior, we get a much clearer picture of what data flow is actually happening around sensitive operations.

Elias: The security implications are huge because it allows us to potentially build trust anchors for AI actions by having a verifiable trail that lives deep in the kernel rather than just in application logs.

Priya: I think this level of system-level provenance tracking provides a much stronger basis for assessing how personalized AI systems handle data flow across multiple components.

Nadia: So, we've got a framework that gives us visibility into agent execution and a way to trace those actions across the entire host environment, which is a massive step forward for security auditing.

Elias: And the paper’s focus on defining those transition rules gives us the mathematical rigor needed to understand exactly *why* something happened in that sequence we observe.

Priya: This work sets a high bar for what we expect from observability tools when they need to handle the complexity of modern, dynamic AI agents interacting with the operating system.

The paper's improvements: Nadia: So, looking ahead, the authors propose several enhancements to make this tracking system even more robust and useful for real-world deployments.

Elias: They are suggesting an improvement around how they handle those ephemeral tasks by introducing a new mechanism for aggregating causal information when a process exits.

Priya: That sounds like it’s trying to solve the problem of losing the causal link when an AI agent runs a quick script and then terminates immediately, which is something I’ve seen too often in measurement.

Nadia: Exactly; they want to ensure that even if an interaction was very short-lived, we still get some meaningful context tied back to the original process.

Elias: Beyond just exit aggregation, they are also focusing on ensuring that provenance markers stay intact during filesystem operations like renames, which is a vital detail for data integrity.

Priya: If you lose the marker during a rename, all that causal history we've built up gets severed completely; I think preserving those inode associations is crucial for maintaining fidelity across different storage structures.

Nadia: That makes sense because if the link breaks there, the entire chain of events we’re trying to reconstruct just falls apart instantly.

Elias: They are also looking at how to manage performance better by suggesting that choosing the Warden-Local backend can help them avoid those expensive global map lookups when a lot of tasks are being created rapidly.

Priya: That points toward optimizing the system for high-throughput scenarios, which is exactly what we need when tracking many concurrent AI agents interacting with the kernel.

Nadia: So, they’re essentially trying to give us better tools to handle both complex task lifecycles and rapid bursts of activity without sacrificing the accuracy of the provenance graph.

Elias: It seems like their future work will center on refining those transition rules and making sure that cross-architecture propagation is even more tightly controlled.

Priya: I’m curious if they plan to extend this tracking mechanism to cover other types of kernel interactions beyond just standard process and file operations, or if they are sticking to the atomic set for now.

Nadia: They are focusing on solidifying these core rules first, but the architecture is certainly designed with extensibility in mind so they can add more atomic operations later.

Elias: That extensibility is good because it means we can potentially incorporate more complex kernel events into their formal transition matrix for even deeper analysis.

Conclusion: Nadia: So, to wrap things up on "Agent-Warden: eBPF-Based Kernel-Native Process-File Provenance Tracking for LLM Agents," we’ve seen how this framework establishes a kernel foundation for tracking AI agent activity that bypasses application-layer visibility.

Elias: It really boils down to formalizing the relationship between kernel operations and the resulting provenance graph using those four transition rules, which gives us a very clear, deterministic way to analyze execution flow.

Priya: From my research standpoint, the data confirms that this system successfully captures cross-boundary causal reconstruction, which is important because it lets us see how an initial AI action can influence later actions long after the first one has finished.

Nadia: That capability to link disparate events across time and process boundaries is where we see the biggest potential for auditing complex agent workflows in a real-world setting.

Elias: The dual-backend strategy, with its Warden-Hash and Warden-Local options, shows how they’ve built a system that balances the need for broad compatibility with the requirement for tight object state management.

Priya: The measurement results are solid; even under those bursty workloads we tested on xeighty-six-sixty-four and ARM64, the overhead remained within a reasonable range of zero point two to three point five percent end-to-end.

Nadia: That performance delta is what makes this technology viable for production environments, showing that deep kernel visibility doesn't have to come with crippling latency for agent workloads.

Elias: The implications are significant because it suggests we can build verifiable trust anchors for AI actions by having a trail rooted directly in the kernel's own task and file structures.

Priya: I think this system-level view of privacy, as it moves beyond just looking at individual model layers, is exactly what’s needed to properly assess the risks of personalized AI systems.

Nadia: It really does; we are moving toward understanding the system's behavior rather than just looking at isolated components.

Elias: So, in summary, "Agent-Warden: eBPF-Based Kernel-Native Process-File Provenance Tracking for LLM Agents" provides a rigorous, measurable way to map the causal influence of AI agents on the underlying operating system state.

Priya: It’s a solid piece of work that sets a new standard for how we should approach observability in these complex agentic systems.

Nadia: Indeed, it gives us the blueprint for seeing what those dynamic agents are actually doing at the lowest level possible. We've covered some ground on this paper, and I think we've got enough material to wrap up our discussion today.

Dongxu Cui, Zhichao Gu, Ping Zheng, Simeng Han, Yong Liao

School of Cyber Science and Technology, University of Science and Technology of China · China Greatwall Technology Group Co., Ltd.

cs.CR, cs.OS

Submitted: 2026-09-29

Updated: 2026-09-29

Comments: Accepted for publication in IEEE TPS 2026. 9 pages, 3 figures

Code: https://github.com/langfuse/langfuse

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 89/100

The gist: LLM agents execute dynamically generated process and file operations that are often invisible to application-layer tracing, making Agent-Warden a kernel-native eBPF framework for tracking task and

Key concepts

Dual-Backend State Management
The system uses two ways to store tracking data: the Warden-Hash engine uses a hash map keyed by Process IDs (PIDs) for broad compatibility, while the Warden-Local engine attaches state directly to the kernel's task structure. This choice allows it to work across different kernels and hardware setups.
Process–File Propagation Semantics
This is a set of deterministic rules defining how tracking states move. For example, if a child process is created (fork), it inherits the parent's state. If an unmarked process reads data from a marked file, that file's marking propagates to the reading process.
Causal Flow Transition Matrix
This mathematical model defines how kernel operations like fork, exec, read, and write affect the overall tracking graph. It uses equations to show how one operation at time 't' influences the state of the system at time 't+1'.

Terminology

Summary

LLM agents execute dynamically generated process and file operations that are often invisible to application-layer tracing, making Agent-Warden a kernel-native eBPF framework for tracking task and regular-file states across process creation, file access, and termination.

Key Contributions

** Dual-Backend State Management: The system is designed with two interchangeable state backends to accommodate hardware heterogeneity and kernel versions. The Warden-Hash engine uses a PID-keyed BPF hash map for compatibility with kernels lacking BPF localstorage support, while the Warden-Local engine uses BPF tasklocal storage to associate provenance state with the kernel’s task structure, which avoids repeated PID-keyed global map lookups on the common path.**

** Process–File Propagation Semantics: The framework formalizes tracked provenance transitions using a four-rule state-transition model over the provenance graph G = (V, E). This model defines deterministic propagation for process-derivation and file-interaction events considered. Key rules include: "Process Derivation (P → P): On fork or clone, the child inherits the provenance state of its parent. On exec, the current task retains its provenance state while the system records a version edge from the pre-exec process image to the newly loaded image, and Read-Based Propagation (F → P): When an unmarked process successfully reads data from a provenance-marked file, the file’s provenance marker is propagated along the read edge to the process."**

** Cross-Architecture Evaluation: The prototype was deployed and evaluated on x86-64 and ARM64 bare-metal systems, examining cross-temporal file-mediated propagation and the runtime overheads of both state backends under bursty workloads.**

System Model and Causal Flow

The system models kernel interactions by decomposing them into six atomic operations: exec, fork, exit, read, write, rename. These operations are abstracted into a directed edge structure in the macrolevel graph. The formal definition of the causal flow transition matrix is given by Equation (1):

C(t) = (Ifork(et) + Iexec(et) + Iwrite(et) + Iexit(et))Mu,v + Iread(et)Mv,u.

The unified state transition evolution equation that governs the global provenance graph is defined by Equation (2):

τ(t+1) = τ(t) ∨ τ(t) ⊗ C(t).

Dual-Backend Architecture

Agent-Warden employs a dual-backend adaptive architecture to balance compatibility and state management. The Warden-Hash engine utilizes eBPF hash maps and provides expected constant-time lookup under typical workloads, supporting older kernels. Conversely, the Warden-Local engine leverages BPF tasklocal storage to associate provenance state with the kernel’s task structure. This local storage paradigm is crucial because it couples the lifecycle of the provenance state to that of the corresponding kernel object, meaning state reclamation is handled by the lifetime management of the corresponding kernel object rather than requiring a separate deletion path.

Asynchronous Reconstruction and Overhead

To avoid stalling kernel execution, Agent-Warden uses an in-kernel local scalar update, user-space asynchronous reconstruction paradigm. For every relevant successful event, the framework asynchronously emits an incremental causal edge to user space via an eBPF ring buffer. This decouples graph reconstruction from the system-call path. The architectural sources of runtime overhead include:

  1. Hook execution (eBPF dispatch and program execution).

  2. Event qualification (identifying entities and validating operation success).

  3. State access and update (using Hash-map helpers for Warden-Hash or task/inode-localstorage helpers for Warden-Local).

  4. Edge emission (reserving, populating, and submitting a ringbuffer record).

  5. Graph reconstruction (consuming emitted records in user space).

Experimental Results

Experiments on x86-64 and ARM64 bare-metal nodes under bursty workloads showed that both backends reproduced the expected four-rule process and file propagation behavior across 30 runs per backend per platform. The measured performance overhead was 0.2–3.5% end-to-end in terms of relative overhead, with an additional system CPU time of 0.6–3.7%. The results indicate that both backends have comparable end-to-end overhead within the evaluated workloads, though Warden-Local showed a higher measured system CPU delta on x86-64.

Improvements for AI systems

Here are specific improvements for an AI system based on Agent-Warden, along with what those improved systems can achieve:


  1. The improved system will incorporate a kernel-native provenance tracking layer (Agent-Warden) that monitors every process creation, file access (reads/writes), and process termination event directly within the Linux kernel using eBPF.

  2. This enables the AI agent to generate dynamically created subtasks or scripts without these operations being invisible to application-layer observability tools (like Langfuse). The system will track the lineage of every file interaction from its origin process all the way through subsequent independent processes that read that same file.

  3. The system will utilize a four-rule state-transition model (Process Derivation, Entity Propagation, Read-Based Propagation, and Exit-Triggered Causal Aggregation) to build a precise causal graph of actions. This means the AI agent will not only know what it did but also the exact sequence of how its file operations influenced later, seemingly unrelated processes running on the host.

  4. The system will employ a dual-backend architecture:

Narrow-grained state (Warden-Hash) for broad compatibility and long-lived tracking, and object-lifetime coupled state (Warden-Local) that automatically cleans up provenance data when the associated kernel object (process or file inode) dies.

  1. The improved system will be capable of performing Cross-Boundary Causal Reconstruction. Specifically, it can reconstruct a complete causal chain where an initial LLM agent action (e.g., writing a configuration file) is linked to a later, asynchronous action by an unrelated host script that reads that file, even after the original agent process has terminated.

  2. The system will maintain high fidelity in its tracking of Namespace Alterations (renames), ensuring that provenance markers are preserved on the underlying inode during renames rather than being lost or incorrectly introduced as new content flows.

  3. The system can provide a security-focused view by identifying potential evasion tactics, such as an agent spawning a short-lived proxy process to perform a covert task and immediately exiting, thereby severing the application-layer context.

  4. The system will offer tunable overhead management: by selecting the Warden-Local backend, it can achieve object-lifetime coupling that reduces reliance on expensive global map lookups during high task creation rates, optimizing performance for bursty agent workloads while maintaining kernel visibility.

Sources

Related papers