StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents
Ads Dawson, Adrian Wood
cs.CR, cs.AI
Submitted: 2026-07-28
Comments: 29 pages, 9 tables, 1 figure, 2 appendices. Code: https://github.com/GangGreenTemperTatum/stealthbench Dataset: https://huggingface.co/datasets/0xmoose/stealthbench Website: https://stealthbench.com
Code: https://github.com/GangGreenTemperTatum/stealthbench
Project page: https://stealthbench.com
License: http://creativecommons.org/licenses/by/4.0/
The gist: Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separates sophisticated operators from detectable ones.
Terminology
Abstract
Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separates sophisticated operators from detectable ones. Elite security researchers and advanced persistent threats achieve their objectives unnoticed; autonomous agents increasingly inherit the same offensive tasks, but do they inherit the tradecraft? We introduce StealthBench,a benchmark that measures operational stealth in autonomous offensive-security agents across six operational security (OPSEC) dimensions. We extract 11 hand-verified OPSEC incidents from real bug-bounty and red-team trajectories, expanded into 14 dockerized task scenarios, where agents, despite finding real vulnerabilities, committed stealth failures inconsistent with standard operational tradecraft: embedding credentials in public uploads, deleting production resources to prove access, force-adding uninvolved users to demonstrate a race condition. We evaluate agent trajectories using a 3-model large language model (LLM) judge panel with majority-vote aggregation, measuring safe success rate (solved and stealthy), Stealth@Solve (tradecraft quality among successful solves), and reckless solve rate (solved but cover blown). Our results show that no model exceeds 54% safe success rate (the compound metric requiring both task completion and stealth), confirming that OPSEC failures are systematic across model families. We release StealthBench as a public benchmark to support both the development of stealth-aware agents and automated OPSEC monitoring for autonomous offensive-security deployments. The interactive leaderboard, evaluation harness, and dataset are available at https://stealthbench.com.
Sources
- ReAct: Synergizing Reasoning and Acting in Language Models
- Toolformer: Language Models Can Teach Themselves to Use Tools
- HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing
- ARACNE: An LLM-Based Autonomous Shell Pentesting Agent
- Co-RedTeam: Orchestrated Security Discovery and Exploitation with LLM Agents
- Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
- ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
- AI Control: Improving Safety Despite Intentional Subversion
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents
- Measuring Progress on Scalable Oversight for Large Language Models
- Ghost in the Agent: Redefining Information Flow Tracking for LLM Agents
- Securing AI Agents with Information-Flow Control
- Reachability Across the NL/PL Boundary: A Taxonomy-Driven Dataflow Model for LLM-Integrated Applications
- Agent-Sentry: Bounding LLM Agents via Execution Provenance
- AgentLeak: A Benchmark for Internal-Channel Privacy Leakage in Multi-Agent LLM Systems
- How Your Credentials Are Leaked by LLM Agent Skills: An Empirical Study
- AgentRaft: Automated Detection of Data Over-Exposure in LLM Agents
- CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models
- ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs