Adversarial Attacks in Multi-Agent LLM Pipelines: Unveiling Structural Vulnerabilities in Agentic AI Architectures
Faisal Haque Bappy, Tahrim Hossain, Tarannum Shaila Zaman, Raiful Hasan, Kamrul Hasan, Tariqul Islam
cs.CR, cs.AI, cs.MA
Submitted: 2026-08-01
Comments: This paper has been accepted at the 2026 IEEE Global Communications Conference (GLOBECOM)
Code: https://github.com/SPaDeS-Lab/adversarial-llm-pipeline
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multi-agent LLM pipelines orchestrate multiple specialized language model agents into structured workflows where intermediate outputs are passed across agents to solve complex tasks.
Terminology
Abstract
Multi-agent LLM pipelines orchestrate multiple specialized language model agents into structured workflows where intermediate outputs are passed across agents to solve complex tasks. This design introduces a security gap absent in single-agent settings: once an agent accepts adversarial content, it is propagated as trusted input throughout the pipeline. We argue that this vulnerability stems from the absence of boundary verification, a security primitive that enforces explicit validation of data as it crosses inter-agent boundaries, including content, identity, execution intent, and state integrity. Without such verification, modern pipelines embed implicit trust assumptions that are not adversarially robust, giving rise to structurally distinct attack surfaces (e.g., content injection, agent impersonation, plan deviation, and memory poisoning). Leveraging annotated production traces from the GAIA and SWE-Bench benchmark, we show that these vulnerabilities arise in benign deployments and largely evade existing evaluation frameworks. We further operationalize these failure modes within a controlled multi-agent setting and evaluate them across GPT-5-mini, Claude Sonnet 4.5, and Kimi K2.5 under identical pipeline configurations. The results reveal that attack success aligns with pipeline structure rather than model capability, indicating that adversarial vulnerability is fundamentally an architectural property and motivating a shift toward pipeline-level defenses.
Sources
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems
- The Dark Side of LLMs: Agent-based Attack Vectors for System-level Compromise
- Open Challenges in Multi-Agent Security: Towards Secure Systems of Interacting AI Agents
- Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
- TRAIL: Trace Reasoning and Agentic Issue Localization
- Audit-LLM: Multi-Agent Collaboration for Log-based Insider Threat Detection
- Flooding Spread of Manipulated Knowledge in LLM-Based Multi-Agent Communities
- AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs