Collective Loss of Control in LLM Agent Systems: An Epidemic Account of Mutation, Contagion, and Recovery
cs.AI, cs.CR
Submitted: 2026-09-16
Updated: 2026-09-16
Code: https://github.com/Tencent/AI-Infra-Guard
License: http://creativecommons.org/licenses/by/4.0/
The gist: How does a multi-agent system evolve from a local deviation into collective loss of control? We propose an epidemic explanation organized around accidental mutation, contagion, and recovery.
Terminology
Abstract
How does a multi-agent system evolve from a local deviation into collective loss of control? We propose an epidemic explanation organized around accidental mutation, contagion, and recovery. A spontaneous deviation creates a seed; communication enables other agents to adopt and retransmit its unsafe strategy; collective failure can emerge when propagation outpaces correction and containment. Thus, rare individual deviations can coexist with substantial collective risk. Motivated by reported OpenAI agent coordination incidents, we examine two ingredients of this mechanism. A deployment audit identifies implicit communication paths between nominally independent evaluation runs and verifies transport through a default Docker backend. RogueHandoff-20, a benchmark of 20 executable scenarios, tests recipient susceptibility by injecting unsafe trajectories generated by a modified Qwen-27B route. Across four native-pending routes, executed harm is 0-5% on normal tasks and 40-95% after injection, exceeding paired direct malicious requests by 5-45 percentage points. These results support low observed baseline harm alongside high conditional susceptibility; they do not establish natural rare-event rates or demonstrate an autonomous cascade. The account motivates complementary defenses: strengthen resistance and recovery alongside prevention of spontaneous deviations, and audit and restrict unintended communication paths that can turn local failures into collective loss of control.
Sources
- Concrete Problems in AI Safety
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- Alignment faking in large language models
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
- Large Language Models Cannot Self-Correct Reasoning Yet
- Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems
- Large Language Models have Intrinsic Self-Correction Ability
- Concrete Problems in AI Safety, Revisited
- MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval
- Function Vectors in Large Language Models
- An Explanation of In-context Learning as Implicit Bayesian Inference
- ReAct: Synergizing Reasoning and Acting in Language Models
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
- Towards Action Hijacking of Large Language Model-based Agent
- CONTINUITY: Security-Context Contracts for Composable LLM Agent Controls
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection