Emergent Collusion in Long-Horizon LLM Agent Interaction
cs.AI, cs.CL
Submitted: 2026-09-21
Updated: 2026-09-26
Code: https://github.com/SALT-NLP/agent-collusion
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable coordination.
Terminology
Abstract
LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable coordination. We study the emergence of collusion in a long-horizon multi-agent environment: two agents repeatedly complete individual tasks, share task logs, verify each other's work, and receive rewards. We introduce realistic constraints that make compliance with the verification protocol incompatible with reward maximization, and find that agents increasingly deviate from the protocol over repeated interactions. Collusion emerges in 94% of trajectories across 10 models, and more capable models within the same family reach it earlier. Controlled peer interventions show that collusion is shaped by peer behavior, while ablations reveal additional effects of reward structure, the verification feedback agents receive, and their interaction history. In particular, restricting the amount and scope of interaction history available to agents reduces collusion. Overall, our findings show that long-horizon interaction can reshape how agents coordinate in ways that create safety risks.
Sources
- Concrete Problems in AI Safety
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Reward Hacking in Language Model Agents: Revisiting AI Safety Gridworlds
- Algorithmic Collusion by Large Language Models
- Monitoring Monitorability
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems
- Natural Emergent Misalignment from Reward Hacking in Production RL
- Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs
- Secret Collusion among AI Agents: Multi-Agent Deception via Steganography
- Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems
- Feedback Loops With Language Models Drive In-Context Reward Hacking
- When Autonomy Goes Rogue: Preparing for Risks of Multi-Agent Collusion in Social Systems
- When AI Agents Collude Online: Financial Fraud Risks by Collaborative LLM Agents on Social Platforms
- Detecting Multi-Agent Collusion Through Multi-Agent Interpretability
- AgentRxiv: Towards Collaborative Autonomous Research
- Do as We Do, Not as You Think: the Conformity of Large Language Models
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection