Intrusion Detection for Agentic Processes: Evidence-Based Runtime Monitoring
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "Intrusion Detection for Agentic Processes".
Elias: Agent deployments increasingly combine language-model inference with retrieval, delegation, tool execution, external-system access, and human approval.
Nadia: First, who's behind it and why it matters.
Title and authors: Nadia: Moving on, let's talk about the title and who wrote this paper, "Intrusion Detection for Agentic Processes: Evidence-Based Runtime Monitoring," because it really sets the stage for what we’re looking at here. It suggests a focus on making security detection work directly within the operational flow of these agents.
Elias: The authors are Arslan Brömme and others, and seeing their background, you can see they've built a product- and vendor-neutral black-box architecture before, so this work feels like an evolution of their earlier ideas into a more focused security layer.
Priya: I wonder what the real impact of this is for privacy researchers; does it give us better ways to measure the data flow within these agents without needing deep access to the core model weights?
Nadia: Well, they propose A-IDS as an evidence-aware security interpretation layer that compares runtime observations against a versioned expectation baseline, which means it’s not just about catching bad things but understanding if what we see matches the established rules.
Elias: That focus on the baseline is where I think the cryptographic and theoretical side gets really interesting; they are treating security policies as a versioned set of expectations with specific conditions for triggering them, which gives us a concrete structure to test against.
Priya: So, it sounds like the core idea is creating a system that can tell us, based on the evidence gathered during execution, whether the agent is following its intended path or if something has gone wrong according to our defined rules.
The paper's summary: Nadia: So, to summarize what this paper lays out regarding "Intrusion Detection for Agentic Processes: Evidence-Based Runtime Monitoring," it’s about proposing A-IDS, which is an evidence-aware layer that evaluates observations against a governed and versioned expectation baseline for things like workflow state and mandatory events.
Elias: The core mechanism they describe involves separating visibility from matching; they treat the process itself as the monitored object and define how observations are captured, normalized, and then compared against those expectations to generate findings.
Priya: What I find interesting is their modeling of evidence claims—they don't just look at raw data; they distinguish between an event, an observation that validates records, a claim about those records, and finally a security finding that compares the claim to the policy.
Nadia: That distinction is crucial because it helps them define what constitutes strong evidence versus weak indicators of trouble, which directly impacts how we interpret any potential issue found during runtime monitoring.
Elias: And they introduce a concept like "source health," which denotes whether an observation source was operational and reliable during the relevant time interval, even though they clarify that this doesn't necessarily establish semantic truth about the data itself.
Priya: That caveat about source health versus semantic truth is important for privacy folks because it acknowledges that we can measure the capture path integrity without claiming absolute knowledge of what happened inside.
The paper's improvements: Nadia: Now, regarding the improvements they suggest for this system, the paper focuses on how A-IDS structures its detection taxonomy by grouping deviations based on the *type* of deviation rather than treating them all as one single category, like process sequence deviation or authorization violations.
Elias: And their modeling of a finding itself is quite detailed; it’s structured as a tuple that includes the window, the expectation, and crucially, an evidence status that summarizes the observation state without implying anything about compromise probability.
Priya: I'm interested in how they handle conflicts; they specifically discuss identifying "Evidence Conflicts" when contradictory observations come from different sources during runtime analysis, which should help operators focus on where the disagreement lies.
Nadia: That conflict detection is a big step because it moves beyond simple matching to highlight areas where the evidence itself is inconsistent, which helps separate technical deviations from their actual security context.
Elias: Furthermore, they address the monitoring-plane attack surface by assuming an attacker doesn't control every relevant observation source and that semantic interpretation should be separated from privileged monitoring functions.
Priya: That separation sounds like a good way to manage risk; if an LLM analyzer can only produce bounded claims without authority over the expectation registry, it limits where prompt injection could have a direct, unchecked impact on the policy baseline.
Conclusion: Nadia: So wrapping up this discussion on "Intrusion Detection for Agentic Processes: Evidence-Based Runtime Monitoring," we see A-IDS offering a structured way to interpret runtime events by comparing them against versioned expectations and clearly classifying deviations based on their type and evidentiary support.
Elias: Essentially, the paper provides a framework for making security interpretation less about guessing what happened and more about systematically checking if the observed evidence satisfies a set of pre-defined governance rules.
Priya: I think the focus on distinguishing between technical deviations and their context, coupled with flagging evidence conflicts, gives us better tools for understanding the actual operational impact of these agentic processes on privacy and data handling.
Nadia: That's what it’s all about; building a system that provides bounded findings that separate evidentiary status from operational impact so we know exactly where to look next when we have an alert.
Elias: And looking ahead, the paper suggests separating semantic interpretation from privileged monitoring functions, which is key for keeping the monitoring plane itself secure against adversarial observation content.
Priya: I think as we move forward, focusing on how these evidence claims are aggregated and what kind of data they reveal about agent behavior will be really important for our field.
cs.CR
Submitted: 2026-09-10
Updated: 2026-09-10
Comments: 7 pages, 3 tables
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 76/100
The gist: Agent deployments increasingly combine language-model inference with retrieval, delegation, tool execution, external-system access, and human approval.
Key concepts
- Observation
- A concrete event in the monitored world. It must bind specific details like actor identity, object, time, and policy version to be useful for security analysis.
- Evidence Claim
- A statement about a property of observations or records. It is not truth itself but a bounded assertion that compares observed data against predefined rules or expectations.
- Governed Expectation
- A versioned set of security policies defining what the system should do, including authorization rules and required events. These expectations are used to match against incoming observations.
- Finding Structure
- The output of the detection system, which links a specific deviation type to evidence status and operational impact. It ensures that minor deviations are not treated with the same severity as potentially destructive actions.
Terminology
Summary
Agent deployments increasingly combine language-model inference with retrieval, delegation, tool execution, external-system access, and human approval. This paper proposes an Agentic-Process Intrusion Detection System (A-IDS), a security interpretation layer for runtime intrusion detection that compares evidence-supported observations against an explicitly governed and versioned expectation baseline for workflow state, authorization, communication, and mandatory events.
The gist
A-IDS defines an evidence-aware security interpretation layer that evaluates dynamically due governed expectations, separates visibility from event matching, represents unresolved evidence explicitly, emits bounded findings that separate evidentiary status from operational impact, and identifies adversarial observation content as a distinct monitoring-plane attack surface.
Model Components: Observations and Evidence Claims
The monitored process is represented through observable events rather than inferred hidden reasoning. The paper distinguishes between an event in the monitored world, a record that technically captures it, an observation that normalizes or validates one or more records, an evidence claim that states a bounded property of those records or observations, and a security finding that compares such claims with governed expectations. A concrete observation must bind at least event type, actor identity, relevant object or counterparty, local time or sequence reference, policy or workflow version, capture source, and an evidence reference or content commitment. The paper notes that A cryptographically committed record provides integrity evidence, not semantic truth or completeness.
Furthermore, a source health denotes whether an observation source was operational and sufficiently reliable as a capture path during the relevant interval,
though this does not establish semantic truth.
Governed Expectations and Matching Logic
For A-IDS, a security policy Pv is treated as a versioned set of governed security expectations and authorization relations, including their scope, applicability conditions, evidence requirements, exception rules, and governance metadata. A concrete expectation e is modeled as a tuple: (type, actor, object, trigger, precedence, deadline, cardinality).
Matching between an observation o and an expectation e is three-valued: true only when the policy-defined type [and] actor [and] object [and] identity [and] temporal [and] cardinality constraints are satisfied.
Visibility is evaluated separately from matching; Vt(e) determines if the required observation paths were Sufficient,
Insufficient,
or Unknown
by checking source health, capture path enablement, and tolerance windows. The due-event match ratio Ct = Mt / Dt, which is described as a narrow matching signal, not a completeness or correctness proof.
Detection Taxonomy and Finding Structure
A-IDS groups initial detection classes by the type of deviation rather than treating them as one homogeneous ontology. Table 2 categorizes deviations into classes such as: Process Unobserved mandatory event, Process Sequence deviation, Authorization Unauthorized interaction, Behavioral escalation, Evidence conflict, Evidence-path anomaly, and Influence. A finding is modeled as: finding = (window, expectation, observation, deviation [type], origin hypothesis (optional), support [binding records/references], evidence status [summarizing observation state], impact).
This structure ensures that Strong evidence for a minor deviation is not equivalent to weak evidence for a potentially destructive action,
and the evidence status must not be interpreted as the probability that an agent is compromised.
Architecture and Monitoring-Plane Security
The novel core of A-IDS is the expectation-and-evidence finding layer, which evaluates governed expectations against observations. Upstream detectors (signature, anomaly, content) are optional evidence producers whose outputs enter A-IDS as bounded claims with their own assumptions.
The architecture separates five functions: Sensors observe agent/human/tool events; Evidence processing authenticates and normalizes records; a Governed expectation registry supplies the baseline; a matching layer evaluates due expectations; and a finding function emits the bounded record. Crucially, A-IDS enforces isolation by ensuring The monitored process should have no authority to modify the monitor’s policy baseline, configuration, credentials, evidence store, or execution state.
This separation is also defined as a trust boundary.
Monitoring-Plane Attacks and Influence Paths
The model considers two distinct attack surfaces: deviations within the monitored agentic process and attacks against the monitoring plane through adversarial observation content. A key concern is that Indirect prompt injection has shown that instructions embedded in data processed by an LLM-integrated application can influence its behavior.
The paper assumes an attacker does not simultaneously control every relevant observation source and the authoritative expectation-definition path, noting that if they did, A-IDS cannot provide meaningful detection assurance.
To address this, semantic interpretation should be separated from privileged monitoring functions; an LLM analyzer can operate as an isolated evidence producer whose output remains a bounded claim, with no authority over the expectation registry.
For influence paths originating from prompt injection, an origin hypothesis requires more than temporal proximity; it needs "an identifiable untrusted source, evidence of exposure or propagation toward the affected actor, and a temporally compatible path to the deviation.
Improvements for AI systems
Here are specific improvements to an existing AI system based on the A-IDS framework proposed in the paper, and what those improved systems could achieve:
-
The system can transition from simple
anomaly detection
(detecting unusual tool calls) to a state-awareEvidence-Aware Security Interpretation Layer.
-
The improved system will monitor not just individual actions, but the entire agentic process graph, comparing runtime events against a versioned, machine-testable security policy baseline.
-
It will explicitly track and report on the state of obligations (e.g., mandatory approvals) using three-valued matching:
-
The system can differentiate between:
-
A
True Match
(the process is secure according to policy), aSufficient Observed
state (a true match exists but visibility is incomplete), anUnobserved
state (no match yet, but the process has been running long enough to rule out immediate non-occurrence), or aViolated
state (a resolved observation directly contradicts the policy). -
The system can generate bounded findings that separate evidentiary status from operational impact:
-
It will distinguish between a technical deviation (e.g., an unauthorized tool call) and its security context, such as whether it is an
Unobserved mandatory event
or aProcess Sequence deviation.
-
The improved system can isolate the monitoring plane itself as an attack surface by treating adversarial content reaching semantic evidence producers as a distinct threat, requiring specific isolation and trust-boundary requirements for those producers.
-
It can implement distributed correlation by explicitly flagging
Evidence Conflicts
when contradictory observations from different sources (e.g., agent self-report vs. gateway log) occur, allowing operators to focus on the conflict rather than just one side of the story. -
The system can perform
Source-Associated Deviation
analysis by linking a deviation back to an untrusted source and a traceable path, providing anOrigin Hypothesis
that distinguishes between mere correlation and potential causal influence (e.g., distinguishing between observation of prompt injection vs. proof of causal control). -
The system can enforce layered defense by separating the monitoring function from the execution environment:
-
It can employ hardened components to capture and normalize raw observations, allowing an isolated LLM-based analyzer to act as a bounded evidence producer without having authority over the expectation registry or enforcement mechanisms, thereby limiting prompt injection risk to that specific analysis layer.
Sources
- A Black Box for Agentic Processes: Blockchain-Anchored Evidence for AI Agent Communication, Human Oversight, and GRC Audits
- An Evidence Model for Agentic Processes: Evidence Claims, Trust Assumptions, and Policy Assessment
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- SentinelAgent: Graph-based Anomaly Detection in Multi-Agent Systems
- VIGIL: Runtime Enforcement of Behavioral Specifications in AI Agent Skills
- Runtime Compliance Verification for AI Agents
- Ghost in the Agent: Redefining Information Flow Tracking for LLM Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs