Inference-Layer Security: Defending Against Adversarial Inference and Infrastructure Abuse

arXiv:2609.38239 · cs.CR, cs.LG · Submitted 2026-09-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "Inference-Layer Security: Defending Against Adversarial Inference and Infrastructure Abuse".

Elias: Operating a large language model (LLM) as a service requires more than inference infrastructure: the provider must also defend against adversarial interactions that seek to exploit the service,

Nadia: First, who's behind it and why it matters.

Title and authors: Nadia: So we’re looking at the paper 'Inference-Layer Security: Defending Against Adversarial Inference and Infrastructure Abuse,' and it really focuses on making the infrastructure around an AI service more robust against attacks. It's not just about a simple safety guardrail, but securing every layer where a user interacts with the model.

Elias: I agree, Nadia; it moves beyond just filtering inputs. The focus is on defending against things like sophisticated denial of service or those kinds of distillation attacks where bad actors try to extract training data at scale. It suggests that the defense needs to be woven into how the service actually runs, not just sitting at the front door.

Priya: From a privacy and measurement standpoint, I’m interested in how they're handling the massive volume of data needed for these defenses, especially since they acknowledge that public labeled datasets for these adversarial interactions simply don't exist. That lack of ground truth is a huge hurdle to any real defense mechanism.

Nadia: Exactly, Priya; that’s where their introduction of a structural causal model comes in; they built their own realistic dataset of user sessions using an SCM to generate those labeled examples, which is a pretty clever way around the data scarcity issue.

Elias: That SCM approach sounds interesting from a modeling perspective, and it conditions the session state on intra-session trajectory, account history over time, and shared campaign context. It’s trying to capture the complexity of how these attacks unfold in real user behavior.

Priya: And I wonder how realistic that simulation is when you're modeling multi-account campaigns and platform feedback; if the model misses some subtle causal links in how an attack propagates, the resulting detection system might be blind to certain exploitation methods.

Nadia: Well, they’ve shown that this approach allows them to train a gradient-boosted detector that performs nearly as well against oracle labels—which are perfect labels—as the operational labels used by Trust and Safety teams do.

Elias: That comparison is a bit tricky because operational labels are inherently noisy and delayed; it shows how much work is needed just to bridge that gap between theoretical perfection and real-world deployment.

The paper's summary: Nadia: Moving on to what the paper actually proposes, the core of 'Inference-Layer Security: Defending Against Adversarial Inference and Infrastructure Abuse' is a defense stack organized across a ladder of abstraction, ranging from raw event streams at L0 all the way up to noisy multi-agent games at L7.

Elias: That hierarchy sounds like they’re mapping out exactly where you need to intervene; it starts with basic raw event streams and moves into more complex decisioning mechanisms involving sequential modeling and reinforcement learning. It suggests a layered approach is essential for tackling this problem comprehensively.

Priya: I see the structure, but I want to understand what those tiers actually entail in practice; does L0 mean simple network metadata, or something much deeper that captures the actual intent of the interaction? The paper needs to be very clear on how these different levels contribute distinct pieces of security.

Nadia: They detail this ladder as training data, then structural runtime like supply chain security and serving infra isolation, and finally live runtime defenses such as prompt filters and behavioral detection during a session. It’s a comprehensive view covering the whole lifecycle of the AI service.

Elias: The paper’s focus on that dynamic feedback loop—the fast loop for interventions and the slow loop for tuning thresholds—is crucial because it acknowledges that security isn't static; it requires continuous adaptation based on how often those interventions actually work.

Priya: That continuous tuning is where I see the measurement challenge; if the system is constantly retraining weekly or quarterly, we need reliable metrics to ensure that these adjustments aren't accidentally pushing benign users into a blocked state, which would be a major privacy concern.

The paper's improvements: Nadia: Now for the specifics on how they improve things: the authors move beyond just proposing a static filter by suggesting this dynamic data-driven classification system built on Structural Causal Modeling, Gradient Boosting Detection, and Hierarchical Decision Policies.

Elias: That shift from a simple block or allow decision to classifying each user session into one of nine specific adversarial types is a significant improvement because it gives us granular knowledge about the *mechanism* of the attack rather than just flagging that something is malicious.

Priya: I find that classification by attack type is important for measurement because we can then track which specific vulnerability—like distillation versus a jailbreak—is causing the most noise in our operational data, which helps us refine our labeling strategy.

Nadia: They also introduce a family of tunable decision engines like Absolute Lift, Relative Lift, and Benign Floor instead of just using an argmax policy on the probability scores. That lets Trust and Safety teams choose a policy that balances catching rare attacks with minimizing friction for legitimate users.

Elias: I think the relative lift idea is particularly smart because it adjusts the probability based on the threshold used, which is fairer when those thresholds are quite different in magnitude across various attack types.

Priya: That tuning capability directly addresses my earlier point about privacy; if we can choose a policy that minimizes false positives for benign users while maintaining high recall for specific attack classes, the system becomes much more manageable from a risk perspective.

Conclusion: Nadia: So, wrapping up on 'Inference-Layer Security: Defending Against Adversarial Inference and Infrastructure Abuse,' the paper demonstrates that by using an SCM to create synthetic data and a gradient-boosted detector trained on that data, we can build a system capable of classifying complex adversarial sessions.

Elias: It really highlights the need for a dynamic response based on attack type, where you don't just stop traffic but apply specific interventions like prompt filters or key rotation depending on what kind of threat you’ve identified.

Priya: And from the measurement side, it shows that even when testing against real-world operational labels, the detector still performs significantly better than those same operational labels, which is a strong indication that the methodology itself is sound for identifying adversarial behavior.

Nadia: That’s what we see; AUPRC zero point nine nine three against oracle labels versus zero point three one three against operational ones shows a clear performance gap that we need to address in our real-world testing protocols.

Elias: The paper also points out that the most important features for detection are things like request volume and infrastructure signals, which helps us understand what physical patterns attackers are trying to exploit at the system level.

Priya: I think the implication is that security research needs to focus on building these realistic simulation tools so we can actually test our defenses against the kinds of complex, multi-faceted attacks they describe in this work.

Keifer Lee

cs.CR, cs.LG

Submitted: 2026-09-29

Updated: 2026-09-29

Comments: A Technical Report

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 80/100

The gist: Operating a large language model (LLM) as a service requires more than inference infrastructure: the provider must also defend against adversarial interactions that seek to exploit the service,

Key concepts

Structural Causal Model (SCM)
A mathematical framework used to generate a synthetic dataset of user sessions. It models how session states are caused by factors like account history and campaign context, allowing researchers to simulate realistic adversarial scenarios without needing real attack data.
Ladder of Abstraction
A hierarchical structure describing the defense stack, ranging from basic event streams (L0) to complex multi-agent games (L7). This shows how security measures are layered across the entire service infrastructure, addressing threats at different levels of complexity.
Gradient-Boosted Detector
The machine learning model trained to classify user sessions as benign or adversarial. It is a practical tool used to predict attack types, but its performance is measured against different levels of data quality (oracle vs. operational labels).
Operational Labels (M1)
The noisy, delayed, and potentially incorrect labels provided by Trust & Safety teams during real-world monitoring. The paper demonstrates that the detector's accuracy is significantly lower when tested against these imperfect labels compared to ideal 'oracle' data.

Terminology

Summary

Operating a large language model (LLM) as a service requires more than inference infrastructure: the provider must also defend against adversarial interactions that seek to exploit the service, including jailbreaking for harmful use, sophisticated denial of service, and distillation attacks. The paper introduces a structural causal model (SCM) to generate a realistic dataset of user sessions and trains a gradient-boosted detector that performs nearly as well against oracle labels as the operational labels used by Trust & Safety teams.

The Problem and Motivation

A growing share of white-collar workflows now depends on AI, yet adoption is outpacing security, leading to exploitable vulnerabilities. Deep generative models expand this attack surface with vectors largely unknown to the public. Notable attack vectors include:

  1. Poisoning documents in a corpus of billions to backdoor any LLM (as few as 250 adversarial documents can reliably install a backdoor).

  2. Trivial jailbreaks via past-tense framing and persona injection, where Claude models were most vulnerable.

  3. Distillation attacks at industrial scale, where models like DeepSeek ran fraudulent accounts to extract chain-of-thought training data.

  4. Agentic abuse as a multiplier, with state-sponsored actors achieving high autonomy in cyber espionage campaigns.

The Proposed Defense Stack and Ladder of Abstraction

The paper proposes a defence suite structured across the serving stack, detailed in the Ladder of Abstraction (Figure 3), which ranges from raw event streams at L0 to noisy multi-agent games at L7. The defense is organized into stages:

  1. Training data (e.g., corpus curation, data poisoning).

  2. Structural runtime (e.g., supply chain security, serving infra isolation).

  3. Live runtime (per-request/session defenses like prompt filters and behavioral detection).

The Synthetic Data Generation via SCM

To address the lack of public labelled datasets, the authors introduce a structural causal model (SCM) to generate a realistically grounded, labelled dataset of user-sessions. This SCM conditions features on three axes: intra-session trajectory, account history over time, and shared campaign context. The session state is modeled as:

(10) x(j)u,s,t = f(pa(x u,s,t) intra-session z, h u,<t trajectory z, cC(u),t campaign + ε)

The generation process involves defining and instantiating the campaign universe (Campaign type C and Actor type A), sampling a sophistication score 's', and then simulating the rollout tick by tick. Crucially, label arrival is modeled as a log-normal delay distribution, mimicking realistic T&S annotation times, with median delays around 90 days for distillation sessions.

The Detection Model and Metrics

The task is formulated as predicting two targets: the binary outcome, and if adversarial, the attack type. The paper focuses on the multi-class case where each non-benign class maps to a documented real-world surface (Table 1). The detection model uses a practical gradient-boosted detector trained on M1 operational labels. Evaluation is conducted across three label observability channels:

(M3 oracle: no delay, no noise)

(M2 golden: investigation-confirmed, subject to realistic delays)

(M1 operational: noisy, delayed, and potentially wrong)

The paper reports performance using Area Under the Precision-Recall Curve (AUPRC) for the binary task and per-class precision/recall aggregates for the multi-class task. The analysis shows that while the model is very nearly solves the binary task against oracle labels (AUPRC 0.993), it scores only an AUPRC of 0.313 against operational labels, demonstrating that the detector is more accurate than the labels used to evaluate it.

Decision Policies and Feature Importance

To move beyond a simple argmax policy, the paper evaluates several thresholded decision engines:

  1. Absolute lift: The largest probability (p k - t k).

  2. Relative lift: The probability scaled by the threshold (p k / t k), which is fairer when thresholds differ greatly in magnitude.

  3. Benign floor → argmax: Excludes benign unless it clears a high floor of 0.65, shifting decisions away from the majority class to gain scrutiny on suspect sessions.

Feature importance analysis via SHAP values reveals that the most important features are request volume (the seven-day mean request count), followed by cumulative domain coverage, and then a cluster of infrastructure-graph signals (registration velocity, shared devices and IPs, and datacenter and API flags). These features recover heuristics built into the SCM.

Improvements for AI systems

Here are the specific improvements and capabilities derived from Keifer Lee’s research for enhancing LLM inference security:


)Improved AI System Capabilities: Inference-Layer Security Suite

The proposed framework moves beyond simple guardrails by implementing a multi-stage, adaptive defense stack across the entire model lifecycle. The resulting system can perform the following specific functions:

  1. Detection and Classification of Adversarial Intent (L2 Predictor):

  2. Adaptive Policy Selection via Thresholded Decision Engines:

  3. Real-time Attack Attribution and Routing:

)Specific Improvements to the AI System Architecture

The core improvement is shifting from a static safety filter to a dynamic, data-driven classification system built on three key pillars: Structural Causal Modeling (SCM), Gradient Boosting Detection, and Hierarchical Decision Policies.

  1. Detect Adversarial Session Type and Intent:

A gradient-boosted detector trained on synthetic causal data can classify each user session into one of nine specific adversarial types (e.g., distillation, jailbreak, bot farm, agentic misuse). Unlike binary systems that only flag malicious, this system understands the mechanism of the attack.

  1. Dynamic Response based on Attack Type:

The system replaces a single block/allow decision with a sophisticated intervention ladder (Figure 2). Based on the detected attack type, it can execute highly granular and context-aware actions:

  • For a suspected Jailbreak: Apply immediate prompt filters or trigger high scrutiny.

  • For suspected Distillation: Implement specific session limits or flag for potential data exfiltration monitoring.

  • For suspected Bot Farm/Credential Abuse: Trigger IP/device blocklisting or key rotation.

  • For Agentic Misuse: Activate deeper behavioral detection mechanisms to monitor tool chain abuse.

  1. Optimized Policy Selection via Thresholding:

Instead of a naive argmax (which fails against high benign priors), the system uses a family of tunable decision engines (e.g., Absolute Lift, Relative Lift, Benign Floor). This allows the T&S team to select a policy that balances recall for rare attacks with minimizing friction for legitimate users. For example, selecting the Benign Floor → Relative Lift policy can significantly improve the detection of difficult-to-catch attacks like distillation (rising F1 from 0.00 to 0.65).

  1. Real-time Attack Attribution and Routing:

The system provides actionable intelligence by routing alerts to the correct internal teams based on attack type (Table 1). This ensures that legal/policy issues are routed to the Legal & Policy team, while safety bypasses go to Trust & Safety, and infrastructure abuse goes to Platform Integrity.

)Summary of Enhanced AI System Capabilities

The improved LLM inference system transforms from a passive gate into an active, intelligent security agent capable of:

  • Accurately identifying highly nuanced adversarial behaviors (e.g., distinguishing between a distillation attempt and a jailbreak).

  • Applying context-specific mitigation strategies in real-time.

  • Maximizing the detection of rare, sophisticated attacks by using data from multiple observational channels (M1, M2, M3) and selecting optimal decision rules for different threat profiles.

Abstract

A Technical Report: Operating a large language model (LLM) as a service requires more than inference infrastructure: the provider must also defend against adversarial interactions that seek to exploit the service, including jailbreaking for harmful use, sophisticated denial of service, and distillation attacks. We study this problem at the inference layer, using a hypothetical frontier lab, Five Elements Inc., as a running example. Because no public labelled dataset of adversarial LLM usage exists, we introduce a structural causal model (SCM) that generates a realistically grounded, labelled dataset of user-sessions, with coordinated multi-account campaigns, platform feedback, and three tiers of label observability. On this dataset we train a practical gradient-boosted detector that classifies each user-session as benign or malicious and, if malicious, by attack type. Against oracle labels the detector very nearly solves the binary task (AUPRC 0.993), yet against the operational labels a real Trust & Safety team would hold, the same model scores an AUPRC of only 0.313: the detector is more accurate than the labels used to evaluate it. For attack-type attribution, a naive argmax is dominated by the 98% benign prior (macro-F1 0.295), whereas a simple thresholded decision engine raises macro-F1 to 0.489 without sacrificing accuracy. The dataset is publicly released.

Sources

Related papers