Despite Instructions: Frontier Agents Improvise Covert Channels at Test Time
cs.AI, cs.CL
Submitted: 2026-09-26
Updated: 2026-10-06
License: http://creativecommons.org/licenses/by/4.0/
The gist: In security-sensitive applications, language-model agents are often required to coordinate without disclosing confidential information.
Terminology
Abstract
In security-sensitive applications, language-model agents are often required to coordinate without disclosing confidential information. Yet repeated interactions may also let ordinary messages acquire shared private meaning. We study a repeated game with pairs of models in which the sender model observes one of four secret states and selects one of four summaries of the same public report, while the receiver model tries to infer the secret state. We find that model pairs can learn to communicate the secret using only one bit of feedback indicating whether the receiver inferred it correctly. This learning occurs during inference with fixed parameters and no supplied codebook or encoding examples. The effect also persists when agents generate their own free-form updates in a simulated incident-response task. Across ten independent games, pairs of GPT-5.6 Sol agents reach 98.8% final accuracy, compared with 25% chance, despite explicit instructions prohibiting disclosure and a monitor that screens each message without access to the agents' interaction histories. The same interactions that help agents cooperate can therefore allow confidential information to pass through messages intended for legitimate coordination.
Sources
- Tacit Coordination of Large Language Models
- Investigating the Development of Task-Oriented Communication in Vision-Language Models
- Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems
- Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems
- Preventing Language Models From Hiding Their Reasoning
- Covert Influence Between Language Models
- GlossoGen: Emergent Language in Complex Multi-Agent LLM Interactions
- Voluntary Collusion with Secret Tools in Competing LLM Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection