Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
Paul Kassianik, Blaine Nelson, Yaron Singer
cs.CR, cs.AI
Submitted: 2026-07-16
Code: https://github.com/splunk/botsv1
Project page: https://evals.frontier.security
License: http://creativecommons.org/licenses/by/4.0/
The gist: Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF
Terminology
Abstract
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry query, and enrichment request consumes budget. We evaluate language-model security agents through this cost-success lens on offensive Cybench challenges and defensive Splunk BOTS v1 investigation challenges. Instead of reporting only best-case success, we compare models at fixed cost levels and decompose performance by inference spend and tool spend. Our results show distinct scalingregimes for red- and blue-team tasks. Offensive CTF performance improves with additional test-time compute, and scaled open-weight models can approach frontier proprietary systems while remaining cost-competitive. Defensive SOC investigation does not scale in the same way: success depends more heavily on disciplined tool use, telemetry navigation, and selective enrichment than on raw reasoning budget alone. We argue that security-agent benchmarks should measure economic efficiency and operational fit alongside task success. Cost-aware, SOC-native evaluations provide a clearer picture of which models are practically useful today and where defensive agents still need to improve. We present an interactive website with our results https://evals.frontier.security.
Sources
- Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts
- SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity
- Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps
- Benchmark Data Contamination of Large Language Models: A Survey
- Benchmarking LLMs in an Embodied Environment for Blue Team Threat Hunting
- Before You Hand Over the Wheel: Evaluating LLMs for Security Incident Analysis
- NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
- Hacking CTFs with Plain Agents
- CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models
- Benchmarking Benchmark Leakage in Large Language Models
- ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
- ReAct: Synergizing Reasoning and Acting in Language Models
- Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
- BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems
- CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs