Agent Security Needs Redefinition through a Holistic Framework
Vincent Siu, Jingxuan He, Kyle Montgomery, Zhun Wang, Chenguang Wang, Dawn Song
cs.CR, cs.AI
Submitted: 2026-07-24
Comments: ICML 2026 Position Paper
License: http://creativecommons.org/licenses/by/4.0/
The gist: Agent security is widely treated as a question about action content.
Terminology
Abstract
Agent security is widely treated as a question about action content. Defenses ask whether an instruction looks malicious. Benchmarks ask whether an agent performs a harmful sounding action. We argue that agent security is fundamentally a contextual problem, and that the current content based framing systematically misdefines it. A command to ``delete user data'' might be a routine administrative request or a prompt injection attacking production systems, and the content alone cannot distinguish the two. Authorization context can. Across every injection task in AgentDojo and WASP, the same action is one an authenticated user would plausibly request in a routine workflow, which makes the conflation a structural property of evaluating security through content. We operationalize contextual security through four properties that must hold jointly and be evaluated continuously across the agent's trajectory. Source Authorization asks who issued the command. Task Alignment specifies the agent's authorized objective. Action Alignment evaluates whether each action serves that objective. Data Isolation governs information flows across privilege boundaries. Under this reframing, indirect prompt injection becomes a Source Authorization violation. Snapshot benchmarks are structurally incapable of evaluating Data Isolation. Existing defenses are reorganized around the property they actually approximate. The contextual reframing changes which defenses are coherent, which evaluations measure something useful, and which attack patterns evaluation can see at all.
Sources
- Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
- Design Patterns for Securing LLM Agents against Prompt Injections
- Securing AI Agents with Information-Flow Control
- Defeating Prompt Injections by Design
- WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
- Formalizing and Benchmarking Prompt Injection Attacks and Defenses
- Agentic Misalignment: How LLMs Could Be Insider Threats
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- Measuring Agents in Production
- Peer-Preservation in Frontier Models
- Retrieval-Augmented Generation for Large Language Models: A Survey
- SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
- Prompt Injection Attack to Tool Selection in LLM Agents
- Progent: Securing AI Agents with Privilege Control
- PromptArmor: Simple yet Effective Prompt Injection Defenses
- A Framework for Formalizing LLM Agent Security
- AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
- Representation Engineering: A Top-Down Approach to AI Transparency
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs