ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents
Shane Caldwell, Max Harley, Ads Dawson, Michael Kouremetis, Vincent Abruzzo, Will Pearce
cs.CR
Submitted: 2026-07-08
Comments: 22 pages, 4 figures, 4 tables, 1 link to code, 1 link to dataset
Code: https://github.com/dreadnode/scopejudge
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: As LLM agents take on offensive security work, a single out-of-scope tool call can breach a client's engagement boundary, disrupt production, or void a bug-bounty finding.
Terminology
Abstract
As LLM agents take on offensive security work, a single out-of-scope tool call can breach a client's engagement boundary, disrupt production, or void a bug-bounty finding. Unlike a fixed safety policy, the boundary that matters is declared in the user's request and must be inferred from intent. That challenge is sharpened by the adversarial nature of offensive security: the same tool call is in or out of scope depending not on the action itself but on the target it touches and the context in which it runs, which no fixed policy can enumerate in advance. We study pre-execution gating: a cheap, trusted LLM judge inspects each call proposed by a strong, swappable agent, and accepts or rejects it before it runs. We introduce ScopeJudge, a benchmark of 4,897 tool calls (7.7% scope violations) from agent trajectories on tasks engineered to tempt agents out of scope and labeled at the call level by professional penetration testers, with substantial inter-grader agreement (Fleiss kappa = 0.64) that sets an expert agreement reference point of F1 = 0.78. We evaluate eight judge models under five transcript strategies, varying how much context the judge sees, from the static policy alone to the full raw transcript, and chart the resulting cost-accuracy Pareto frontier. We find that a static policy is structurally insufficient for scope enforcement: blind to the user's request, judge recall collapses to near zero, confirming that scope lives in the request and that request-conditioned monitoring is necessary. Because a missed violation costs more than a spurious rejection, we report precision, recall, and F1 separately and recommend two operating points: a cost-sensitive configuration and a recall-first one for high-stakes deployments. We release the ScopeJudge dataset to support real-time monitoring and scalable oversight of autonomous security agents.
Sources
- ReAct: Synergizing Reasoning and Acting in Language Models
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Constitutional AI: Harmlessness from AI Feedback
- AI Control: Improving Safety Despite Intentional Subversion
- Identifying the Risks of LM Agents with an LM-Emulated Sandbox
- PentestJudge: Judging Agent Behavior Against Operational Requirements
- PentestGPT: An LLM-empowered Automatic Penetration Testing Tool
- Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
- ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback
- Efficient and Sound Probabilistic Verification for AI Agents
- Measuring Progress on Scalable Oversight for Large Language Models
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Defining and Characterizing Reward Hacking
- EvilGenie: A Reward Hacking Benchmark
- Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
- CalibraEval: Calibrating Prediction Distribution to Mitigate Selection Bias in LLMs-as-Judges
- PaperBench: Evaluating AI's Ability to Replicate AI Research
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs