MAS-Shield: A Defense Framework for Secure and Efficient LLM MAS
cs.MA, cs.AI, cs.CR
Submitted: 2025-11-28
Updated: 2026-09-17
Comments: EMNLP findings 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large Language Model (LLM)-based Multi-Agent Systems (MAS) are susceptible to linguistic attacks that can trigger cascading failures across the network.
Terminology
Abstract
Large Language Model (LLM)-based Multi-Agent Systems (MAS) are susceptible to linguistic attacks that can trigger cascading failures across the network. Existing defenses face a fundamental dilemma: lightweight single-auditor methods are prone to single points of failure, while robust committee-based approaches incur prohibitive computational costs in multi-turn interactions. To address this challenge, we propose MAS-Shield, a secure and efficient defense framework designed with a coarse-to-fine filtering pipeline. Rather than applying uniform scrutiny, MAS-Shield dynamically allocates defense resources through a three-stage protocol: (1) Critical Agent Selection strategically targets high-influence nodes to narrow the defense surface; (2) Light Auditing employs lightweight sentry models to rapidly filter the majority of benign cases; and (3) Global Consensus Auditing escalates only suspicious or ambiguous signals to a heavyweight committee for definitive arbitration. This hierarchical design effectively optimizes the security-efficiency trade-off. Experiments demonstrate that MAS-Shield achieves a 92.5% recovery rate against diverse adversarial scenarios and reduces defense latency by over 70% compared to existing methods.
Sources
- Combating Adversarial Attacks with Multi-Agent Debate
- Training Verifiers to Solve Math Word Problems
- Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
- Large Language Model based Multi-Agents: A Survey of Progress and Challenges
- Red-Teaming LLM Multi-Agent Systems via Communication Attacks
- Measuring Massive Multitask Language Understanding
- Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems
- More Agents Is All You Need
- Secure Multi-LLM Agentic AI and Agentification for Edge General Intelligence by Zero-Trust: A Survey
- Large Language Model Agent: A Survey on Methodology, Applications and Challenges
- Self-critiquing models for assisting human evaluators
- Metacognitive Self-Correction for Multi-Agent System via Prototype-Guided Next-Execution Reconstruction
- Large Language Model Safety: A Holistic Survey
- CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge
- PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety
- An Electoral Approach to Diversify LLM-based Multi-Agent Collective Decision-Making
- Demonstrations of Integrity Attacks in Multi-Agent Systems
- MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration
- JudgeLM: Fine-tuned Large Language Models are Scalable Judges
Related papers
- Highway Congestion Reduction through Reinforcement Learning Based Eulerian Headway Control
- You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents
- Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems
- MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimization
- PeroMAS: A Multi-agent System of Perovskite Material Discovery
- StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning