Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents
Harry Owiredu-Ashley
cs.CR, cs.AI, cs.CL
Submitted: 2026-07-08
Comments: 8 pages, 6 figures. Code and artifacts: https://github.com/Harry-Ashley/action-graded-severity
Code: https://github.com/Harry-Ashley/action-graded-severity
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Ignore Previous Prompt: Attack Techniques For Language Models
- Agent-SafetyBench: Evaluating the Safety of LLM Agents
- Defending Against Indirect Prompt Injection Attacks With Spotlighting
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs