Non-Degenerate Risk Certification for Automated Security Decisions: A Decision-Contract Theory with ATT&CK-Aligned Triage as a Worked Instance

arXiv:2608.12444 · cs.CR, cs.LG · Submitted 2026-08-12 · Read on arXiv

Zhenpeng Li

Guangzhou Health Science College

cs.CR, cs.LG

Submitted: 2026-08-12

Updated: 2026-08-14

Comments: 17 pages, 2 figures, 10 tables

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 100/100

The gist: Author: Zhenpeng Li (Guangzhou Health Science College) arXiv:2608.12444v1 [cs.CR] 12 Aug 2026 --- The paper addresses a fundamental flaw in unconditional risk bounds for automated security decisions.

Terminology

Summary

Author: Zhenpeng Li (Guangzhou Health Science College)

arXiv:2608.12444v1 [cs.CR] 12 Aug 2026


The paper addresses a fundamental flaw in unconditional risk bounds for automated security decisions. The authors state: An unconditional risk bound on automated decisions can be satisfied without automating anything, since a selector that never acts drives the bound to zero. They observed this failure mode directly: We initially observed exactly this failure mode in one training run of an ML+CRC baseline, where a LightGBM classifier satisfied its FAR target by never automating a single alert. The authors emphasize this is structural: This is not a bug in one classifier's calibration; it is a structural consequence of what an unconditional risk bound can and cannot certify.

The motivating context is Security Operations Centers (SOCs) processing thousands of intrusion detection system (IDS) alerts daily, where attributing each alert to a MITRE ATT&CK technique is critical. The authors note: Manual attribution is accurate but does not scale: the median SOC analyst handles 20–30 alerts per hour [1], while enterprise IDS deployments generate orders of magnitude more. They propose using large language models (LLMs) for automated attribution but warn: a misattributed alert is worse than an unattributed one: it can trigger incorrect containment procedures, waste analyst time on the wrong investigation track, or mask a genuine threat behind a benign-looking attribution.

The paper formalizes a decision contract as a pair C = (g, E) where g is a selector determining whether the system acts and E is a reflexive acceptance relation defining correctness. The authors prove an error-conservation law (Theorem 3): B(h) = R(C)+D(C)+M (C) where B(h) is the base classifier's fine-grained error, R(C) is harmful automated risk, D(C) is deferred error, and M(C) is semantic masking. The paper states: a fixed base classifier's error B(h) does not disappear under a risk certificate, it is only reassigned among harmful automation, human deferral, and semantic masking.

Theorem 4 gives an exact fine-to-coarse risk transfer identity: R φ(g) = R fine(g) − M φ(g), where M φ(g) is the within-fiber confusion mass. Theorem 5 establishes the impossibility of reverse risk transfer: if any fiber of φ contains two distinct labels, then R φ(g) = 0 does not imply any non-trivial bound on R fine(g).

The paper derives an exact geometric characterization (Proposition 2): "Γ τ(x) = 1 ⇐⇒ p(2)(x) < τ ≤ p(1)(x)." This leads to the singleton capacity (Theorem 6): κ(f) = sup τ [F2−(τ) − F1−(τ)], which "is the maximum action rate attainable by any global singleton threshold. If κ(f) < ρ, no calibration method using a single global threshold can satisfy an action-rate requirement A ≥ ρ, regardless of α or calibration-set size."

The paper also introduces the risk-feasible capacity (Definition 4): κ α(f) = sup A(τ): τ ∈ [0,1], R(τ) ≤ α, which requires labels but correctly separates threshold misalignment from a risk-constrained limit that κ(f) alone cannot see.

Definition 5 defines an (α, ρ)-actionable contract requiring R(C) ≤ α and A(C) ≥ ρ. Theorem 8 shows this excludes the all-abstain solution by construction: e.g., α = 0.05, ρ = 0.80 jointly certify at least 75% correct automation and at most 6.25% conditional error among automated outputs.

Proposition 5 provides a finite-sample action-rate certificate using a one-sided Hoeffding bound: A(τ̂) ≥ A δ:= max 0, Â(τ̂) − ε m(δ) with probability at least 1−δ. The authors note: "The two sides of the certificate carry genuinely different guarantees... R(C τ̂) ≤ α is the standard split-conformal marginal statement of Theorem 1... A(τ̂) ≥ A δ is a genuine finite-sample, (1−δ)-confidence lower bound for the fixed, already-realized τ̂."

The paper evaluates across:

  • 3 IDS datasets: CIC-IDS-2018, HIKARI-2021, RT-IoT2022

  • 6 LLMs: Gemma-2 9B, LLaMA-3 8B, Mistral 7B, Qwen-3 8B, Qwen-3 14B, Qwen-3 32B

  • 4 error-rate thresholds: α ∈ 0.01, 0.05, 0.10, 0.20

  • 2 ML baselines: XGBoost and LightGBM with the same CRC abstention layer

The ATT&CK mapping is a "deterministic bijective mapping φ: C attack → T from attack categories to ATT&CK techniques" (Table I), mapping DoS→T1498, CredentialAccess→T1110, Exploitation→T1190, Probe→T1046.

Across 18 configurations, the mean utility is 0.834, meaning that 83.4% of attack alerts receive a correct automated technique attribution. Under the stringent criterion of empirical FAR ≤ α across all 5 random seeds, 65 of 72 configurations pass (90.3%). The 7 exceedances are concentrated in four model–dataset pairs.

The mean automation rate, which also includes wrong singleton attributions, is 0.850. The three low-utility outliers are: CIC × Mistral (0.378), CIC × Gemma (0.516), and RT-IoT × Gemma (0.063). The best configurations on HIKARI achieve 98.5% utility with FAR = 0.

  • P1, P2, P4, P7 (accounting identities): "All four hold with zero violations across the 18-configuration, 4-α matrix (difference < 10−9 throughout)."

  • P5 (semantic masking): Using real fine-grained attack subtypes from CIC-IDS-2018 (7 DoS subtypes, 4 CredentialAccess subtypes, 3 Exploitation subtypes), the identity R φ(g) = R fine(g) − M φ(g) "holds exactly at all 4 α levels (difference < 10−9 in every case). The realized masking mass is small: M φ(g) = 2.1 × 10−5 at α = 0.05."

  • P3 (capacity diagnosis): Every configuration in the LLM matrix has κ̂ ≥ 0.80, ruling out structural incapacity at any deployment floor ρ ≤ 0.80. However, κ̂ alone would misclassify two of the three as recoverable threshold misalignment; κ̂ α shows only CIC × Gemma-2 actually is. Specifically: RT-IoT × Gemma-2 (κ̂=0.803, κ̂ α=0.122, A(τ̂)=0.111) is risk-constrained incapacity; CIC × Mistral (κ̂=1.000, κ̂ α=0.380, A(τ̂)=0.380) is risk-constrained incapacity; CIC × Gemma-2 (κ̂=0.915, κ̂ α=0.637, A(τ̂)=0.552) is threshold misalignment, confirmed by exhibiting an alternative threshold reaching A=0.605 at empirical risk 0.036 ≤ α.

  • P6 (actionability certificate): Certifying (α, ρ)-actionability at α = 0.05, ρ = 0.5 against A δ (not the point estimate) still passes 16 of 18 configurations.

  • CRC vs. Argmax: B1-Argmax violates the FAR target in 6 of 18 cases... while B3-CRC satisfies it in 16 of 18.

  • ML+CRC: On all three datasets, ML+CRC is competitive with the best LLM+CRC (utility within 1–3 percentage points).

  • Retracted structural-incapacity claim: "Retraining LightGBM on the identical data... with 20 different random seeds... produced a degenerate (Utility < 0.1) model in 0 of 20 runs. The original collapse is therefore best explained as a one-off training artifact—plausibly floating-point nondeterminism in multi-threaded histogram construction."

FAR, utility, and coverage are numerically identical at both levels across all seeds and α because the mapping is bijective.

"Empirical FAR remains below α in all 8 configurations despite exchangeability not being guaranteed across datasets—a robustness observation, not a formal transfer guarantee—but utility degrades substantially in several cases, down to 4.3%."

Utility rises sharply from n cal = 50 (0.895) to 100 (0.954), then saturates within 0.01 through 800... 200 labeled attack samples suffice for near-optimal CRC deployment.

The paper proposes a four-stage deployment procedure:

  1. Capacity check: "before calibration, estimate κ̂ on an unlabeled readiness sample (Proposition 3). If κ̂ falls below the required action-rate floor ρ, no calibration choice can reach ρ (Theorem 6); replace the base classifier rather than proceeding."

  2. Calibration: fix τ̂ on a labeled calibration split via Algorithm 1, certifying R(C τ̂) ≤ α (Theorem 1).

  3. Actionability certification: on a disjoint certification split, compute the action-rate lower bound A δ (Proposition 5). Only if A δ ≥ ρ does the contract satisfy (α, ρ)-actionability with confidence 1 − δ.

  4. Deployment: "alerts with g τ̂(x) = 1 are logged with the singleton ATT&CK technique and routed to automated playbooks; alerts with g τ̂(x) = 0 are forwarded to analysts."

The paper acknowledges: "Our LLM+CRC evaluation uses 3–4 ATT&CK techniques per dataset, whereas the full ATT&CK matrix contains over 600 techniques. Also: The controlled FAR is an unconditional risk over all attack alerts. It should not be read as the conditional error rate among the subset of alerts that receive automated attributions. The exchangeability assumption is noted: In practice, network traffic is non-stationary... Adaptive conformal methods that reweight or refresh calibration data over time are a natural extension."

The paper concludes: a risk certificate is a property of a decision contract, not of a predictor alone, and can be satisfied by weakening that contract through abstention or semantic coarsening rather than genuine predictive competence. The practical implication: a security automation pipeline should certify capacity, risk, and action rate jointly, not risk alone.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems and what the improved systems can do:

  • Improvement: Modify the risk-certification layer to require joint certification of both risk (R ≤ α) and action rate (A ≥ ρ), rather than risk alone. Implement the (α, ρ)-actionable contract check using the finite-sample Hoeffding lower bound on A.

  • Capability: The AI system can no longer game its risk certificate by abstaining on all inputs. It must demonstrate genuine automation utility, preventing degenerate all-abstain solutions that satisfy risk bounds vacuously.

  • Improvement: Add a capacity-check stage before any calibration or training, estimating the singleton capacity κ̂ on an unlabeled readiness sample. If κ̂ < ρ (the required action floor), automatically flag the base classifier as structurally incapable and trigger replacement.

  • Capability: The system can detect threshold misalignment versus fundamental incapacity before wasting compute on calibration. It avoids deploying classifiers that can never meet action-rate requirements regardless of tuning.

  • Improvement: Compute both κ̂ (unconstrained capacity) and κ̂ α (risk-feasible capacity) during evaluation. Use the gap between them to classify failure modes: if κ̂ ≥ ρ but κ̂ α < ρ, report risk-constrained incapacity (requires better base model); if κ̂ α ≥ ρ but current A(τ̂) < ρ, report threshold misalignment (fixable by re-calibration).

  • Capability: The system gives actionable diagnostic feedback rather than a single pass/fail. Operators know whether to retrain, re-calibrate, or replace the model.

  • Improvement: When the AI maps fine-grained classes to coarse categories (e.g., attack subtypes → ATT&CK techniques), compute the within-fiber confusion mass M φ(g) and verify the identity R φ(g) = R fine(g) − M φ(g) holds empirically. Flag deployments where M φ(g) is non-negligible.

  • Capability: The system can quantify how much semantic masking hides fine-grained errors behind coarse-label correctness. This prevents false confidence when coarse labels look accurate but fine-grained errors are being concealed.

  • Improvement: Track both empirical FAR and utility (correct automation rate) in production, not just FAR. Set alert thresholds when utility drops below a floor while FAR remains acceptable, indicating the model is degrading in a way that risk-only monitoring misses.

  • Capability: The system detects silent performance degradation that pure risk monitoring would miss—e.g., when a model starts automating only easy cases or masking errors through coarsening.

  • Improvement: Use the saturation curve (utility rising sharply from n cal=50 to 100, then flat through 800) to recommend minimum calibration sizes. For new deployments, require n cal ≥ 200 labeled samples for near-optimal performance.

  • Capability: The system avoids under-calibrated deployments that would show artificially low utility, and avoids wasting labeling budget beyond the saturation point.

  • Improvement: When deploying across non-exchangeable data distributions, run the full certification pipeline on the target dataset and report utility degradation explicitly. Do not claim formal guarantees; instead, report empirical FAR compliance as a robustness observation.

  • Capability: The system provides honest uncertainty about cross-domain performance, preventing overconfident deployment in shifted environments while still flagging when FAR holds empirically.

  • Improvement: For any coarse-label prediction (e.g., ATT&CK technique), maintain a secondary fine-grained error tracker that records within-fiber misclassifications. Report both coarse FAR and fine-grained error rate to stakeholders.

  • Capability: The system surfaces hidden errors that coarse-label metrics obscure, enabling better risk communication and more informed human oversight.

The improved system can:

  • Certify genuine automation value—it cannot claim risk compliance while doing nothing

  • Diagnose failure modes precisely—distinguishing retrain the model from re-tune the threshold from replace the architecture

  • Quantify hidden semantic errors—showing where coarse labels mask fine-grained mistakes

  • Provide honest cross-domain performance—with explicit caveats, not false guarantees

  • Guide resource allocation—knowing exactly how many labeled samples are needed and when to stop labeling

  • Detect silent degradation—catching utility collapse even when risk metrics remain green

  • Prevent wasted compute—screening out structurally incapable models before expensive calibration

These improvements transform the AI from a system that merely satisfies a risk bound into one that demonstrably delivers useful, safe, and honest automation.

Sources

Related papers