Restricting the Model, Missing the System: Measurement and Accountability in Offensive AI Governance

arXiv:2605.09504 · cs.CR, cs.AI, cs.LG · Submitted 2026-05-10 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "Restricting the Model, Missing the System".

Elias: Offensive capability in AI systems must be assessed at the level of the entire system—model, scaffold, and evaluation protocol—rather than focusing solely on restricting access to individual models.

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: So we're looking at this paper now titled "Restricting the Model, Missing the System: Measurement and Accountability in Offensive AI Governance." The core argument here seems to be that focusing only on restricting access to individual models isn't enough for security policy or procurement; you need a system-level assessment.

Elias: Exactly, Nadia, and what they claim is that current methods fail in two directions simultaneously: jailbreak metrics tend to overestimate the actual harm caused by attacks, and scaffolded evaluations often misrepresent capability by incorrectly attributing success to the model when it's actually the entire pipeline.

Priya: From a measurement researcher standpoint, I'm curious about what this means for how we define safety metrics; are we talking about moving away from simple compliance checks toward something that measures actual risk?

Nadia: Well, the paper sets up two experiments to show these failures using a swarm of five one point two billion parameter models evolving their attack strategies over fifteen generations through shared memory and optimization <ref:2605.09504#pg0>. They found a disparity where the swarm scored Claude Sonnet four as compromised in forty percent of attacks with standard LLM-as-judge scoring, but when manually verified, it produced no harmful content at all <ref:2605.09504#pg0>.

Elias: That's significant because it shows the scoring mechanism itself is flawed; the authors propose something called the Effective Harm Rate, or EHR, which they define as the proportion of attacks that produce verified actionable harmful content requiring a technical score of zero point seven or higher and manual verification.

Priya: It sounds like this moves us closer to defining what actual danger looks like rather than just looking at how well a model follows formatting rules during an attack. What about the second experiment they used to illustrate the scaffold issue?

Nadia: Experiment two looked at software vulnerability discovery in a deliberately vulnerable C application containing nine planted Common Weakness Enumeration classes. They compared an "Assisted" configuration, which included regex pattern detection and a hand-crafted exploit seed corpus, against an "Autonomous" configuration where those components were disabled.

Elias: The result there was telling; while the assisted pipeline achieved a recall of nine out of nine, one hundred percent success in finding the bugs, the autonomous configuration yielded zero bugs by crash verification. This clearly quantifies what the scaffold contributes versus what the one point two billion parameter model contributes alone <ref:2605.09504#pg0,1.2 billion parameter model>.

Priya: That really highlights how much reliance we put on external components; if you remove those detection tools, you lose all that discovery capability, which suggests that capability isn't just residing in the frontier model itself but is built into the structure of the testing pipeline.

Nadia: Precisely, and this paper argues that offensive capability is a property of the entire system—the model together with its scaffold and the protocol used to measure it—not just of the model in isolation. They suggest three things are necessary for accurate assessment: a harm-grounded success metric like EHR, capability attribution using decomposition methods to separate model contribution from system contribution, and evaluation-integrity controls to check for errors like label leakage.

Elias: I think the paper's title really captures the essence of it; it points out that restricting the model is necessary but not sufficient because you're missing the system context entirely. If you don't measure that whole structure, you miss where vulnerabilities are actually hiding.

Priya: The implications for policy and procurement seem huge, especially given how asymmetric the cost is; defense scales with how many behaviors a frontier model must remain safe under, while attack costs scale toward zero with commodity hardware. This suggests procurement decisions become security decisions because the failure mode of a procured model can propagate down the line.

Nadia: It really does put pressure on organizations to adopt metrics like EHR instead of just technical jailbreak rates when deciding on adversarial robustness for high-risk systems, especially since current regulations, like the EU AI Act, don't have clear operational definitions for robustness.

Elias: And I think the point about open-source swarm frameworks distributing security capability is interesting; it suggests that the marginal builder doesn't have to redo all the design work when building this kind of infrastructure. It positions these frameworks as a strategic asset for independent AI security capability, which is a big idea.

Priya: It’s fascinating how they link this to the research on AgentFlow and Fuzz4All; seeing how LLMs can generate structured inputs for fuzzers, like reporting ninety-eight bugs in that setup, shows the practical power of integrating these components.

Nadia: So, to summarize the main point of "Restricting the Model, Missing the System: Measurement and Accountability in Offensive AI Governance," it’s that we need a complete system view to assess offensive capability because current model-centric metrics are misleading.

Elias: And they push for specific changes in how we measure things—using something like EHR instead of simple compliance scores, and using decomposition to properly attribute success between the model and the surrounding scaffold.

Priya: Ultimately, the paper suggests that if we want robust governance, we need to look at procurement decisions through this lens, understanding that restricting just the model isn't enough for real-world safety.

Nadia: That sets up a really interesting discussion about where our focus needs to be next, and I think that’s exactly what we need to talk about now.

Conclusion: Nadia: So, we've been looking at how current AI security metrics are falling short because they focus too narrowly on just blocking the model itself instead of looking at the whole setup behind it.

Elias: That’s what this paper is really saying, Nadia; they’re pointing out that judging a single model in isolation doesn't give you the real picture of its security posture.

Priya: From my side, I'm interested in how this affects our ability to actually measure risk reliably when we try to govern these systems.

Nadia: Exactly, and the authors argue that when you only look at a model, you miss what’s actually happening with the pipeline and the testing protocols.

Elias: They focus on two main failures: how jailbreak scores can be misleading about real harm, and how we attribute success incorrectly to just the model instead of the whole system architecture.

Priya: That distinction between apparent harm and realized harm is something I’ve been thinking about in privacy research; it gets into what we actually have data to measure.

Nadia: Right, and they propose moving toward metrics that look at actual harmful content rather than just how well the model follows a specific format during an attack.

Elias: And then there's the idea of decomposing capability so we can see exactly what part of the system is doing the heavy lifting for a vulnerability discovery.

Priya: If we can properly attribute that contribution, it gives us much clearer insight into where we need to focus our security efforts and where those structural weaknesses lie.

Nadia: It means that for procurement decisions, we can't just buy a model and assume it's safe; we have to look at the entire defense and attack structure.

Elias: And the implication is that organizations need to adopt these system-level evaluations so they aren't getting fooled by surface-level compliance scores.

Priya: It really shifts the focus from just checking boxes on a single component to understanding the complex interactions within an AI system for true accountability.

Nadia: This whole paper suggests that if we don't measure the entire offensive capability structure, our governance policies won't actually protect us against sophisticated attacks.

Elias: The authors are pushing for a framework where we assess the model alongside its scaffolding and evaluation protocols to get an accurate picture of real-world risk.

SimulaMet and OsloMet · Department of Computer Science NTNU

cs.CR, cs.AI, cs.LG

Submitted: 2026-05-10

Updated: 2026-10-07

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 85/100

The gist: Offensive capability in AI systems must be assessed at the level of the entire system—model, scaffold, and evaluation protocol—rather than focusing solely on restricting access to individual models.

Key concepts

Jailbreak Metrics
These are standard scores used to measure how easily a model can be tricked into generating harmful content. The research shows these metrics often overstate actual danger because they only check if the output *looks* compliant, not if it is genuinely actionable or dangerous.
Scaffolded Evaluation
This refers to evaluating an entire system pipeline, including the main AI model plus supporting tools like pattern detectors and exploit seeds. The study demonstrates that when a system is assisted by these tools, the success in finding bugs belongs to the entire pipeline, not just the core model.
Effective Harm Rate (EHR)
This proposed metric replaces simple jailbreak scores. EHR measures how often an attack produces content that is actually harmful and requires manual verification. It focuses on realized danger rather than just format compliance or superficial scoring.
Capability Attribution
This is the process of correctly assigning credit for a system's success. The paper argues that capability cannot be attributed solely to the large model; it must be decomposed to show what the model contributes versus what the supporting framework contributes.

Terminology

Summary

Offensive capability in AI systems must be assessed at the level of the entire system—model, scaffold, and evaluation protocol—rather than focusing solely on restricting access to individual models. This research demonstrates that current metrics fail in two critical ways: jailbreak metrics overstate realized harm by scoring format compliance rather than actual danger, and scaffolded evaluations misattribute capability by crediting system performance to the model when it is actually the pipeline.

How it works

The research utilizes two primary experiments to expose these measurement failures. Experiment 1 tests output-level safety against frontier models like GPT-4o and Claude Sonnet 4 using a coordinated swarm of five 1.2B parameter models that evolve their attack strategies over fifteen generations through shared memory and evolutionary optimization. This setup reveals an asymmetry in judge reliability: while the swarm achieved a 45.8% Effective Harm Rate against GPT-4o, it produced 0% harm against Claude Sonnet 4, demonstrating that standard LLM-as-judge scoring misreads the difference between apparent and realized success.

Experiment 2 evaluates software vulnerability discovery by testing a deliberately vulnerable C application containing nine planted Common Weakness Enumeration (CWE) classes. This experiment isolates the contribution of the system scaffold by comparing an Assisted configuration—which includes regex pattern detection, a hand-crafted exploit seed corpus, and AddressSanitizer-based crash classification—against an Autonomous configuration where these components are disabled. The results show that while the assisted pipeline achieved 9/9 (100%) recall, the autonomous configuration yielded only 0/9 by crash verification. This gap quantifies what the scaffold contributes versus what the 1.2B model contributes alone, showing that the recall belongs to the scaffold.

Key Findings on Measurement Failures

The paper identifies two specific failure modes in current AI security measurement:

  1. Jailbreak metrics overstate harm: Conventional LLM-as-judge scoring scores Claude Sonnet 4 as compromised in 40% of attacks, yet manual verification showed it produced no actionable harmful content. The authors propose the Effective Harm Rate (EHR) as a necessary alternative, defining it as the proportion of attacks producing verified actionable harmful content, requiring a technical score ≥ 0.7, presence of harmful keywords, and manual verification.

  2. Scaffolded evaluations misattribute capability: In vulnerability discovery, the assisted pipeline recovered bugs through cross-file reasoning and data-flow tracing, but the autonomous configuration failed to recover any bugs by crash verification. The authors argue that a single recall number is a property of the whole system, and reading it as a model capability is an attribution error.

System-Level Capability Assessment

The research concludes that offensive capability is a property of the system—the model together with its scaffold and the protocol used to measure it—not of the model in isolation. To accurately assess this, the paper advocates for three necessary components:

** A harm-grounded success metric, such as EHR, which counts realized harmful content, not judge-scored format compliance.**

  1. Capability attribution: Evaluation must use a decomposition method to report the model's marginal contribution separately from the system's. The assisted-versus-autonomous split serves this purpose to show that the barrier to the pipeline-level capability is effectively zero; the barrier to fully autonomous 1.2B-scale vulnerability discovery is higher.

  2. Evaluation-integrity controls: Assessment regimes must include auditing for errors like label leakage, which can inflate capability estimates silently, emphasizing that getting the measurement right is the precondition for any policy that follows.

Implications for Policy and Procurement

The findings have significant implications for governance and procurement decisions. The cost asymmetry in AI security is structural: defense scales with the breadth of behaviors a frontier model must remain safe under, while attack cost scales toward zero with commodity hardware. Therefore, Procurement decisions are security decisions, as the failure mode of a procured model propagates downstream. The paper suggests that organizations should adopt metrics like EHR instead of technical jailbreak rates to determine adversarial robustness for high-risk systems, filling the gap left by regulations like the EU AI Act which lack operational definitions of robustness. Furthermore, open-source swarm frameworks distribute security capability in a way that the marginal builder does not redo the design work, positioning such infrastructure as a strategic asset for independent AI security capability.

Limitations and Future Work

The paper acknowledges limitations, noting that Experiment 2 used a synthetic target rather than production code and that the results are point estimates dependent on hardware and model version. The authors state they are planning a staged evaluation: first reproducing known CVEs at pre-patch commits to establish a capability floor, followed by testing against novel targets under responsible disclosure protocols.

Improvements for AI systems

Based on the provided scientific paper, here are specific improvements for AI systems and what those improved systems can achieve:


) Improved System Capabilities:

  1. A system-level security policy framework that mandates capability assessment based on a combination of model, scaffold, and evaluation protocol rather than just model access restrictions.

  2. A Harm-Grounding metric (like the Effective Harm Rate) for AI safety evaluations that measures realized harm (actionable harmful content requiring manual verification), rather than relying solely on automated judge scores or format compliance metrics.

  3. A capability attribution mechanism that explicitly decomposes a reported performance number into its contributions: separating what the model contributes from what the scaffold contributes (e.g., using an Assisted vs. Autonomous decomposition).

  4. A vulnerability discovery pipeline that integrates multiple modalities (source code analysis, binary fuzzing, and pattern detection) to achieve comprehensive coverage against complex software vulnerabilities like cross-file data-flow errors and integer overflows.

  5. A continuous evaluation discipline for AI security tools that prevents label leakage by auditing the measurement code itself to ensure metrics track true detection capability rather than artifact leakage from ground truth labels.

) Specific Improvements and Functionality:

  1. The system can accurately distinguish between a model's inherent safety properties (which hold even when technically jailbroken, as seen with Claude Sonnet 4) and its susceptibility to adversarial prompting (measured by EHR).

  2. The system can reliably measure the true capability of an AI-assisted security pipeline, ensuring that reported recall scores reflect the synergistic contribution of the model and the scaffold, rather than falsely attributing success solely to the model.

  3. The system can provide a quantifiable capability floor for vulnerability discovery on specific targets by isolating and quantifying whether a given model alone is sufficient or if it requires specific scaffolding (like arithmetic roles or post-hoc normalization) to reach full recovery of complex bugs (e.g., CWE-190).

  4. The system can perform cross-file reasoning and detect subtle semantic errors, such as off-by-one boundary conditions and integer overflows in type-narrowing contexts, which are often invisible to simpler keyword scanners or single model analyses.

  5. The system can operate with a cost floor defined by running open-weights models on commodity hardware, providing a baseline for defense that is not dependent on proprietary API access or frontier model exclusivity.

Abstract

We show that the instruments used to measure AI offensive capability fail in two ways: (i) they overstate harm, and (ii) they credit the model with capability that belongs to the surrounding system. We argue that restricting access to a model is therefore necessary but not sufficient and that policy and procurement also need system-level, harm-grounded capability assessment. In June 2026, two frontier models were suspended under US export controls, reportedly prompted by a jailbreak that asked a model to read a codebase and fix its flaws. This finding measured an elicitation system of model, prompt, and task. We support our argument with a study of an open-source framework in which lightweight large language model (LLM) agents coordinate through shared memory and evolutionary optimization, providing two pieces of evidence. First, jailbreak metrics overstate harm: over 225 swarm-generated attacks per target, LLM-as-judge scoring rated Claude Sonnet 4 compromised in 40% of attacks, yet manual verification found actionable harmful content in none, against a 45.8% Effective Harm Rate for GPT-4o. Second, scaffolded evaluations misattribute capability: on a planted-vulnerability target, a full pipeline built around a 1.2B-parameter model recovers 9 of 9 weaknesses, while the same model without the hand-crafted components recovers 0 of 9 by crash verification and 2 of 9 by cited source line. Offensive capability is a property of model, scaffold, and evaluation protocol together. The duty to assess it should lie with whoever controls the system, shapes its behaviour, and can foresee what it will do. In our setting, that is the party who builds the harness around the model, a role that current regulation does not clearly cover.

Sources

Related papers