Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation".
Jane: The paper was written by Yuchen Ling, Shengcheng Yu, Zhenyu Chen and Chunrong Fang from State Key Laboratory for Novel Software Technology, Nanjing University, China and Technical University of Munich, Germany.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Paper Discussion Segment 1: Tom: Now that we've established the scope using the full title, let's move into what the paper actually summarizes. When we talk about "Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation," it’s not enough just to list threats; we need to understand how those threats propagate through an agent’s internal processes.
Jane: The key insight they provide in their summary is that the risk isn't centralized in one spot—it's distributed across the entire interaction loop. They move us past simply worrying about what the user types into the prompt and start looking at what happens *after* the model decides to execute a tool or modify its memory.
Lu: The lifecycle view is paramount here; they detail how information flows from an initial prompt, through a planning stage, to tool execution, and then back into memory. Every one of those handoffs is a potential vulnerability point that needs rigorous examination.
Meng: What I found particularly useful in the summary is how they emphasize the need to differentiate between model limitations and systemic architectural flaws. Sometimes we mistake poor performance for a fundamental security failure, but this paper helps us separate those concerns.
Lalam: It provides this necessary taxonomy of risk—it tells us that a flaw in the *system* (like tool access control) is different from a flaw in the *model* itself, and both must be addressed simultaneously.
Tom: So, the summary essentially gives us a mental checklist of every major component an agent uses: input validation, planning modules, external APIs...
Jane: And for each one, they point out specific failure modes that haven't received enough architectural attention yet. For instance, the interaction between external knowledge retrieval and the core reasoning model is complex and often under-controlled.
Lu: The rigor here is that they aren't just pointing fingers at weaknesses; they are providing a structured way to model how those weaknesses interact with each other in a cascading failure scenario.
Meng: It gives us the language to talk about agent failure modes that go beyond simple hallucinations and involve actual capability misuse.
Lalam: This systemic understanding is what elevates the discussion from academic curiosity to an actionable engineering requirement for building trustworthy AI.
Paper Discussion Segment 2: Tom: We’ve seen the summary, and it’s clear that the paper provided a massive expansion of our understanding of agent risk. But what does this research actually suggest we need to *improve* in our current development practices? That's where the practical implications lie for "Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation."
Jane: The biggest takeaway is that individual defensive components are insufficient. They encourage us to move beyond isolated security patches and instead focus on building a cohesive, integrated defensive stack that can handle multiple failure modes simultaneously.
Lu: I agree entirely; the concept of "compositionality" is what they emphasize—the defenses must be designed so they fit together seamlessly, rather than being just a pile of unconnected guardrails. We need reliable stacks.
Meng: From an engineering standpoint, this means we can't just slap on input validation and assume that solves everything. We have to model the interaction between the layers: how does the runtime monitor handle an input that bypasses the initial validator?
Lalam: It forces us toward a truly holistic security view where every component—the memory, the tool, the prompt—is treated as a piece of one integrated whole that must coordinate its trust boundaries.
Tom: It’s fascinating because they don't just give us general advice; they provide detailed guidance on how these defenses should be measured and evaluated.
Jane: This moves us toward needing better governance mechanisms, which is perhaps the most underdeveloped area in the industry right now—how do we govern an agent that is constantly learning and evolving?
Lu: The lack of convergence on a single, universally effective defensive architecture remains the primary challenge for the next phase of development. We are too fragmented right now.
Meng: I keep returning to the practical tradeoff: how do we increase security dramatically without crippling the operational utility or making the agent unusable in a real-
Paper discussion segment 3: Tom: We’ve covered the findings of "Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation," but let's talk about what this research suggests we need to improve in our current development practices moving forward.
Jane: The authors make a huge case that current security measures are like isolated building blocks; they’re useful individually but not designed to stack up or work together reliably. They don're calling for a cohesive, compositional defense architecture instead of just patches on the way.
Lu: I think that’s where the theoretical magic is—moving from focusing on prompt-level risk to studying how risks propagate through a full lifecycle is a monumental shift in mindset, it fundamentally changes how we view system integrity.
Meng: If we're talking about practical implementation, it implies that our systems need more than just one guardrail; they must be designed with explicit trust boundaries at every single operational handoff point to function robustly.
Lalam: It also has a massive cultural implication for how we approach the design of autonomous AI, forcing us to build systems where verifiable authority and state provenance are core values, not afterthoughts.
Tom: So, while prompt injection is visible today, the real focus of this paper is on addressing those deeper systemic concerns like persistent state corruption and multi-agent risks.
Jane: It shows that we have to be far more deliberate about how we model that a security failure isn't just one piece breaking down in time. The contamination can survive and reappear when the next agent or the next planning step kicks in.
Lu: That stateful nature of the risk is what makes this so much harder to simply solve, because it demands a level of systemic assurance that we haven're currently missing entirely.
Meng: When we look at multi-agent propagation, it means that if one agent gets compromised, the risk can cascade through coordination channels to completely unrelated parts of another system. We have to ensure our architectures can handle that whole sequence.
Lalam: This suggests that our future designs must account for how information flows across multiple entities, not just how a single model responds to a prompt or execute a single tool command.
Tom: It’s clear from "Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation" that the complexity of networked failures is now the real story for us. This whole conversation has really highlighted the need for better governance in our AI development process.
Conclusion: Tom: We've covered the deep dive into "Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation," and it's truly a comprehensive look at where our field is right now. It’s clear this research presents a robust framework for thinking about agent security not as an isolated problem but as an entire system.
Jane: Exactly; the paper successfully argues that securing LLM agents requires us to move beyond just prompt safety and start thinking about true secure agent engineering instead of relying on isolated model patching.
Lu: I think the clarity that the state and information flow are inseparable from our future designs is a huge conceptual win for my research.
Meng: The need for better tool governance is also a practical reality we can't ignore, demanding that our systems build hard boundaries around execution capabilities.
Lalam: We must apply this framework to shift our mindset toward creating an AI culture where trust and state provenance are treated as first-class engineering concerns.
Tom: That's the big picture—moving beyond mere functionality to a commitment to security architecture is essential for the long-term success of these technologies.
Jane: It’s a great way to wrap up this discussion, realizing that secure LLM agents require explicit trust boundaries and principled privilege control for the coming years.
Lu: It’s a very inspiring call toward seeing how these vulnerabilities propagate through the entire lifecycle of multi-agent systems.
Meng: We've certainly learned our lessons on where to focus our practical defense efforts, especially when looking at tool-mediated risks.
Lalam: And it’s also a great reminder that this entire field needs to be driven by state-aware, secure AI culture rather than just focusing on model outputs.
Yuchen Ling, Shengcheng Yu, Zhenyu Chen, Chunrong Fang
State Key Laboratory for Novel Software Technology, Nanjing University, China · Technical University of Munich, Germany
cs.CR, cs.AI
Submitted: 2026-08-23
Updated: 2026-08-25
Code: https://github.com/openclaw/openclaw
Project page: https://ling-yuchen.github.io/LLMAgentSecuritySurvey
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 92/100
The gist: Large language model (LLM) agents are rapidly moving from conversational interfaces to software components that plan, invoke tools, maintain memory, and act on external environments.
Key concepts
- Agent Lifecycle View
- This concept views an agent's operation as a flow of information through distinct stages: initial prompt, planning, tool execution, and memory updates. Security must be examined at every handoff point in this process.
- Compositional Defense Architecture
- Instead of using isolated security patches or guardrails, the paper advocates for building a cohesive defense stack. Defenses must fit together seamlessly to handle multiple failure modes simultaneously.
Terminology
Summary
Large language model (LLM) agents are rapidly moving from conversational interfaces to software components that plan, invoke tools, maintain memory, and act on external environments. This transition fundamentally changes the nature of security risk; in agentic settings, failures are no longer limited to unsafe text generation.
Instead of merely problematic text, failures can manifest as hijacked workflows, unauthorized tool use, corrupt[ing] persistent state [or leaking] sensitive information, or trigger[ing] harmful external actions.
The paper addresses the current fragmentation in LLM agent security research—which is fragmented across attack families, defense layers, application domains, and evaluation settings
—by synthesizing a curated corpus of 247 papers. The authors utilize a lifecycle-based, systems-oriented framework that models agent security around the interaction of information flow, delegated authority, and persistent state.
The study organizes the literature around four key research questions: how LLM agent security should be modeled; which threat surfaces and attack families dominate; what defenses have been proposed and with what tradeoffs; and how security claims are evaluated.
Key Findings:
-
Dominant Threats: The analysis finds that
prompt injection and tool-mediated control-flow hijacking still dominate the field.
However, the literature is evolving, withpersistent state corruption and multi-agent propagation becoming central emerging concerns.
-
Defense Limitations: Current defenses are found to be
weakly compositional,
meaning they do not form a coherent, reusable stack. -
Evaluation Gaps: Existing benchmarks are insufficient because they
underrepresent long-horizon, stateful, and deployment-sensitive risks.
The paper argues that securing LLM agents requires a fundamental shift in approach: secure LLM agents require explicit trust boundaries, principled privilege control, provenance-aware state management, and evaluation practices aligned with realistic operational settings.
Methodological Framework:
The authors propose a lifecycle model that connects various stages of the agentic loop—inputs, planning, decisions, tool execution, outputs, memory/state and coordination
—to show how risks emerge and propagate across inputs... to output.
This allows for a unified view where prompt injection, memory corruption, tool-mediated abuse, and multi-agent risk
can be connected within a shared account of the agentic loop.
Conclusion:
The paper concludes that LLM agent security must be treated as a systems problem—not merely an extension of prompt-level model safety. The findings suggest that the field is moving from single-step prompt compromise toward longer-horizon and networked forms of failure,
indicating a need for "a shift in mindset: from securing model outputs in isolation to engineering governable agentic systems whose authority, memory, and coordination structure are treated as first-class security design concerns in deployment."
Improvements for AI systems
Based on a rigorous analysis of this survey, we are not simply patching vulnerabilities; we are fundamentally re-architecting the security posture of LLM Agents. The core issue identified is that existing systems treat agentic failure as isolated input errors, when in reality, risk propagates across the entire operational loop (Input to Planning to Decision to Tool Execution to State).
The following improvements translate the findings into specific engineering mandates.
- Implementation of Lifecycle-Based Trust Boundaries:
We will transition from a simple input filter
model to a multi-layered, lifecycle-aware security stack (A = I, P, D, T, M, O, C to). We will explicitly define and enforce trust boundaries at the interface between every stage of the agentic loop.
- Mandatory Provenance-Aware State Management:
All persistent state (M) will be stored with explicit metadata detailing its origin (e.g., User Prompt,
Tool Output,
Internal Planning
). This allows us to track the source of every piece of memory and implement trust decay—mechanisms that automatically quarantine or expire state items if they are deemed potentially compromised or outdated, preventing long-horizon contamination.
- Enforcement of Capability Scoping (Least Privilege):
Tool invocation (T) will no longer be treated as a single action. We will implement fine-grained access control for every tool, defining specific, verifiable permissions that must be satisfied before the agent is allowed to execute. This prevents malicious tool mediation
by ensuring the agent cannot misuse a tool outside its defined scope (e.g., a read file
tool cannot be coerced into executing system commands).
- Formalized Multi-Agent Protocol and Topology Containment:
We will replace simple communication channels (C) with a formal Role-Scoped Authority framework. Every message passed between agents must carry verifiable cryptographic provenance (who sent it, what role they assume) to prevent coordination failures
or message-level spread.
This allows us to implement topology-aware containment, isolating compromised agent subgraphs.
- Integrated Runtime Guardrails and Policy Enforcement:
We will deploy a dedicated runtime monitoring layer that operates after the initial plan but before execution (D to T). This guardrail system will enforce policies (e.g., Do not call external API X if the input was determined to be from a Web Content source that has been flagged
) and utilize context purification to strip malicious intent from planning or tool-selection steps.
The improved AI system will possess the following capabilities:
-
Resilient Execution: It will successfully execute complex, multi-step tasks even when faced with subtle, indirect prompt injections (e.g., hidden instructions in retrieved web content), because the system will refuse to treat low-authority content as a control signal.
-
Self-Healing and Degradation: If contamination occurs in the persistent state (M), the system can identify the poisoned entry via provenance tracking, quarantine it, and proceed with a verified plan, preventing
delayed or recurring compromise.
-
Secure Delegation: The system can delegate tasks to other agents without risking catastrophic failure. If one agent is compromised (e.g., through a malicious tool output), the system's operational boundaries prevent that compromise from spreading to other parts of the workflow or corrupting shared state.
-
Assurance-Driven Operation: The system will be designed not just to
work,
but to remain governably useful. It can operate under real-world constraints (cost, latency) while continuously providing auditable evidence that its decisions adhere to defined security policies and authorized capabilities.
Abstract
Large language model (LLM) agents are rapidly moving from conversational interfaces to software components that plan, invoke tools, maintain memory, and act on external environments. This transition changes the nature of security risk. In agentic settings, failures are no longer limited to unsafe text generation. Untrusted content may redirect control flow, misuse tool privileges, corrupt persistent state, leak sensitive information, or trigger harmful external actions. At the same time, research on LLM agent security is expanding quickly but remains fragmented across attack families, defense layers, application domains, and evaluation settings. This paper synthesizes 247 papers through a lifecycle-based, systems-oriented framework that models agent security around the interaction of information flow, delegated authority, and persistent state. We organize the literature around four questions: how LLM agent security should be modeled, which threat surfaces and attack families dominate, what defenses have been proposed and with what tradeoffs, and how security claims are evaluated. We find that prompt injection and tool-mediated control-flow hijacking still dominate the field, while persistent state corruption and multi-agent propagation are becoming central emerging concerns. We further find that current defenses provide useful building blocks but remain weakly compositional, and that existing benchmarks still underrepresent long-horizon, stateful, and deployment-sensitive risks. We argue that secure LLM agents require explicit trust boundaries, principled privilege control, provenance-aware state management, and evaluation practices aligned with realistic operational settings.
Sources
- LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection Challenge
- Firewalls to Secure Dynamic LLM Agentic Networks
- Human Society-Inspired Approaches to Agentic AI Security: The 4C Framework
- Semantic Intent Fragmentation: A Single-Shot Compositional Attack on Multi-Agent AI Pipelines
- ACIArena: Toward Unified Evaluation for Agent Cascading Injection
- Exposing Weak Links in Multi-Agent Systems under Adversarial Prompting
- Information-Theoretic Privacy Control for Sequential Multi-Agent LLM Systems
- Securing Generative AI Agentic Workflows: Risks, Mitigation, and a Proposed Firewall Architecture
- Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
- AgenTRIM: Tool Risk Mitigation for Agentic AI
- Design Patterns for Securing LLM Agents against Prompt Injections
- DoomArena: A framework for Testing AI Agents Against Evolving Security Threats
- AgentBound: Securing Execution Boundaries of AI Agents
- VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents
- ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents
- Agents at Risk: How Users Unwittingly Undermine LLM Safety
- AgentGuard: Repurposing Agentic Orchestrator for Safety Evaluation of Tool Orchestration
- SafeMind: Benchmarking and Mitigating Safety Risks in Embodied LLM Agents
- Evaluating the Robustness of Multimodal Agents Against Active Environmental Injection Attacks
- LlamaFirewall: An open source guardrail system for building secure AI agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs