The Ethics of Autonomous AI Agents for Offensive Security
Andreas Happe, Jürgen Cito, Jasmin Wachter
cs.CR, cs.AI
Submitted: 2026-07-22
Comments: accepted at FAIEMA 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLM-driven autonomous agents are reshaping offensive security.
Terminology
Abstract
LLM-driven autonomous agents are reshaping offensive security. Unlike traditional penetration-testing tooling - deterministic, narrowly scoped, and operated by trained practitioners - agentic security tools exhibit indeterminacy along three independent dimensions. First, their actions are drawn from a non-deterministic policy whose outputs resist both ex-ante and ex-post explanation. This complicates incident attribution and pre-deployment safety reviews. Second, their impact is open-ended due to their non-deterministic actions, agency of utilized models, and opaque LLM supply-chains. Third, their user population is indeterminate in both size and required skill: the operating skill floor for using or developing offensive capabilities has dropped sharply. These three properties are linked thematically, but are not derivable from one another. Combined with the structural cost asymmetry between offense and defense, they enable the industrialization of offensive capability. The net short-term effect favors attackers, even if the same technology may, in the long run, democratize access to defensive practice. Existing dual-use cybersecurity and AI-ethics frameworks struggle to address this combination. Our work analyzes how moral attribution becomes diffuse between users, tool-makers, and third parties when employing autonomous AI agents for offensive security. We also examine the stakeholder impact of this technology and provide stratified recommendations.
Sources
- Cochise: A Reference Harness for Autonomous Penetration Testing
- Recognition Without Mitigation: Ethical Frameworks in Autonomous Offensive-LLM Agent Research
- VulnBot: Autonomous Penetration Testing for A Multi-Agent Collaborative Framework
- CAI Fluency: A Framework for Cybersecurity AI Fluency
- Criminal Liability of Generative Artificial Intelligence Providers for User-Generated Child Sexual Abuse Material
- Towards Reliable Local Security Agents: Verifiable Post-Training for Linux Privilege Escalation
- Enhancing Linux Privilege Escalation Attack Capabilities of Local LLM Agents
- Structured access: an emerging paradigm for safe AI deployment
- Incalmo: An Autonomous LLM-assisted System for Red Teaming Multi-Host Networks
- AutoPT: How Far Are We from the End2End Automated Web Penetration Testing?
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs