Attack Success Rate Is Not a Number: On Measurement Validity in Agentic AI Security Evaluation
cs.CR
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: 7 pages, 2 tables, 2 Figures
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Attack success rate (ASR) is the headline metric in nearly every published evaluation of attacks on, and defenses for, LLM agents.
Terminology
Abstract
Attack success rate (ASR) is the headline metric in nearly every published evaluation of attacks on, and defenses for, LLM agents. We argue that ASR as currently used is not a single quantity but a family of metrics parameterized by six design choices that papers seldom specify and never hold constant across the literature. We support this with two studies that require no proprietary access. First, a full-text meta-analysis of 259 agentic-security papers posted to arXiv between February 2025 and September 2026 finds that most report neither a variance estimate nor repeated runs for their headline attack metric: 58% (95% CI 44-71) in a hand-coded random sample of 50, 65.3% by automated coding of all 259. Only 30.9% disclose enough about decoding to establish whether their evaluation was even stochastic, and of the 64 papers we confirm use an LLM judge, 29.7% report any agreement check against human labels. Second, an analytical study shows that these omissions are not cosmetic: on a 100-instance benchmark, the minimum difference in ASR detectable at conventional power is 18.2 percentage points, and two defenses whose true ASRs differ by 5 points are ranked in the wrong order by a single-run evaluation roughly 21% of the time. Because several of the six axes shift ASR in a system-dependent way, the resulting incomparability is not a constant offset that cancels in comparison. We conclude that cross-paper ASR comparison is currently unsupported, and propose a ten-item reporting checklist targeted at each failure we measure. Our aim is not to dispute any individual result but to supply the shared measurement contract the field has so far done without.
Sources
- Ignore Previous Prompt: Attack Techniques For Language Models
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
- Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
- SoK: The Attack Surface of Agentic AI - Tools and Autonomy
- The Attack and Defense Landscape of Agentic AI: A Comprehensive Survey
- The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis
- Prompt Injection Attacks on Agentic Coding Assistants: A Systematic Analysis of Vulnerabilities in Skills, Tools, and Protocol Ecosystems
- Model Context Protocol (MCP) at First Glance: Studying the Security and Maintainability of MCP Servers
- Rethinking MCP Security: A Large-Scale Study of Runtime MCP Servers and Security Scanner Reliability
- Same Name, Different Server: A Security Census of Silent Drift in the Model Context Protocol Ecosystem
- Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry
- REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems
- SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning
- SeClaw: Spec-Driven Security Task Synthesis for Evaluating Autonomous Agents
- Defeating Prompt Injections by Design
- Jailbroken: How Does LLM Safety Training Fail?
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs