On Understanding, Identifying, and Mitigating Vulnerabilities in Agentic Large Language Models

arXiv:2608.10530 · cs.CR, cs.AI · Submitted 2026-08-11 · Read on arXiv

Md Jafrin Hossain, Mohammad Arif Hossain, Nirwan Ansari

Florida International University · Middle Tennessee State University · New Jersey Institute of Technology

cs.CR, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 81/100

The gist: This systematic literature review, conducted under PRISMA 2020 guidelines, examines security vulnerabilities in agentic Large Language Models (LLMs).

Terminology

Summary

This systematic literature review, conducted under PRISMA 2020 guidelines, examines security vulnerabilities in agentic Large Language Models (LLMs). The authors screened 743 records across six databases and retained 85 papers published between 2023 and 2025. The review finds that Attack research outpaces defense work by 3.9:1, with perception-layer vulnerabilities (prompt injection, jailbreaking, adversarial perturbations) dominating the literature, accounting for 66% of papers, while action-layer vulnerabilities (tool misuse, code injection, sandbox escape) appear in only 4.7% of papers, which is misaligned with real-world risk. Code execution security accounts for 3.5% of papers, and tool-augmented agents 12%.

The paper contributes a four-layer vulnerability taxonomy mapping 13 vulnerability types across perception, brain, action, and interaction layers. The perception layer includes direct prompt injection, indirect prompt injection, jailbreaking attacks, and adversarial input perturbations. The brain layer includes backdoor attacks, reasoning manipulation, goal/plan hijacking, and memory poisoning. The action layer includes tool manipulation, function hijacking, code injection, and sandbox escape. The interaction layer includes inter-agent message injection, agent impersonation, and knowledge base corruption.

The authors identify seven open problems centered on containment: (1) security of code-execution agents (only 3 papers), (2) security of embodied agents (0 papers), (3) security of single-agent systems (0 papers primarily focused), (4) maturity of detection methods (10 detection methods vs 52 attacks, a 5:1 ratio), (5) coverage of tool-augmented agents (12%), (6) real-world deployment (most research is lab-based), and (7) standardized evaluation (no unified protocols despite 8 benchmark datasets).

The review's quantitative analysis reveals a 3.9:1 attack-to-defense ratio and a fourteen-fold disparity between perception-layer and action-layer research coverage. Multi-agent systems are the most represented architecture (46% of papers), followed by RAG-based agents (29%), tool-augmented agents (12%), and web agents (4.7%). The paper concludes that Agentic LLM insecurity stems from architectural coupling, where weak isolation allows vulnerabilities to propagate across layers.

Improvements for AI systems

Improvements to AI Systems:

  1. Implement Layer-Aware Isolation Architecture
  • Redesign agentic LLM systems with strict compartmentalization between perception, brain, action, and interaction layers.

  • Use hardware-enforced sandboxes for code execution, separate memory stores with access control, and cryptographic signing for inter-agent messages.

  • Resulting capability: A system where a prompt injection in the perception layer cannot propagate to tool execution or memory poisoning—containment of attacks by design.

  1. Add Proactive Action-Layer Defense Mechanisms
  • Integrate runtime monitors that validate every tool call, code snippet, and file operation against a whitelist of allowed actions.

  • Deploy anomaly detection on agent behavior (e.g., unexpected file writes, network calls, or privilege escalations) with automatic rollback to last safe state.

  • Resulting capability: An agent that can safely execute arbitrary code in a sandbox while detecting and neutralizing tool misuse or sandbox escape attempts in real time.

  1. Develop Cross-Layer Attack Detection Ensemble
  • Combine perception-layer detectors (e.g., prompt injection classifiers) with brain-layer reasoning monitors (e.g., consistency checks on chain-of-thought) and action-layer audit logs.

  • Use a central security orchestrator that correlates signals across layers to identify multi-stage attacks.

  • Resulting capability: Detection of complex attacks that span layers, such as an indirect prompt injection that manipulates memory and then triggers a malicious tool call—something current single-layer detectors miss.

  1. Create Standardized Adversarial Benchmark for Agentic Systems
  • Build a unified evaluation suite with 52 attack types mapped to the four-layer taxonomy, including 10 defense baselines.

  • Include realistic multi-agent, RAG, and tool-augmented scenarios with both lab and simulated production environments.

  • Resulting capability: A measurable, reproducible way to compare agent security across architectures, enabling rapid iteration on defenses and identifying gaps (e.g., zero coverage for embodied agents).

  1. Introduce Memory Integrity Verification
  • Add tamper-evident logging and hash-chaining to agent memory stores.

  • Periodically validate memory contents against expected state using a separate verification module that cannot be influenced by the agent’s own reasoning.

  • Resulting capability: Prevention of memory poisoning attacks where an attacker injects false facts or instructions that persist across sessions, ensuring long-term reliability.

  1. Build Self-Containing Agent Runtime
  • Equip agents with a containment reflex: when a security violation is detected at any layer, the agent automatically terminates the current task, revokes all tool permissions, and enters a read-only diagnostic mode.

  • Include a human-in-the-loop escalation protocol for ambiguous cases.

  • Resulting capability: An agent that fails safe under attack, preventing cascading failures in multi-agent systems or production deployments.

  1. Add Interaction-Layer Authentication for Multi-Agent Systems
  • Implement mutual TLS with per-agent certificates and message signing for all inter-agent communications.

  • Add a central registry that tracks agent identities, roles, and allowed communication patterns.

  • Resulting capability: Prevention of agent impersonation and inter-agent message injection, which currently affect 46% of studied architectures.

  1. Create a Security-Aware Planning Module
  • Modify the brain layer’s planning to include a security cost function—evaluating each proposed action for risk (e.g., executing untrusted code, accessing sensitive data) before execution.

  • Use a separate, non-LLM policy engine to enforce hard constraints on high-risk actions.

  • Resulting capability: An agent that proactively avoids dangerous actions (e.g., refusing to execute code from untrusted sources) rather than relying solely on post-hoc detection.

  1. Develop Adaptive Defense Fine-Tuning
  • Use the 52 documented attack patterns to generate adversarial training data for the agent’s perception and brain layers.

  • Periodically retrain the agent with new attack samples from the literature, focusing on the 3.9:1 attack-to-defense gap.

  • Resulting capability: An agent that becomes progressively more resistant to known attack families, with measurable improvement on the standardized benchmark.

  1. Add Real-World Deployment Safety Wrappers
  • For production agents, include a shadow mode that runs all actions in a simulated environment first, comparing expected vs. actual outcomes before committing to real-world effects.

  • Log all actions with full provenance for post-incident forensic analysis.

  • Resulting capability: Safe deployment of agents in high-stakes domains (e.g., finance, healthcare) where a single misstep could cause significant harm, with full auditability.

Abstract

Large Language Models (LLMs) have undergone a shift from stateless conversational interfaces to autonomous agents capable of multi-step planning, tool invocation, code execution, and maintaining persistent memory. When these agents operate with real-world privileges---calling APIs, modifying files, and querying databases---a compromised reasoning step can trigger unauthorized data access, irreversible state changes, or cascading failures, yet the security research community has not kept pace. To quantify the state of the field, we conducted a systematic literature review under PRISMA 2020 guidelines across six databases, screening 743 records and retaining 85 papers (2023--2025) on agentic LLM security. Attack research outpaces defense work by 3.9:1. Perception-layer vulnerabilities (prompt injection, jailbreaking, adversarial perturbations) dominate, accounting for 66% of papers, while action-layer vulnerabilities (tool misuse, code injection, sandbox escape) appear in only 4.7%, misaligned with real-world risk. Code execution security accounts for 3.5%, and tool-augmented agents 12%. We contribute a four-layer taxonomy mapping 13 vulnerability types across perception, brain, action, and interaction layers, and identify seven open problems centered on containment. Agentic LLM insecurity stems from architectural coupling, where weak isolation allows vulnerabilities to propagate across layers.

Sources

Related papers