Why LLM Agents Collapse Without Oversight: The Enforcement Gap as the Mechanism Behind Emergence World Failures
cs.AI
Submitted: 2026-09-14
Updated: 2026-09-23
Comments: 27 pages, 3 figures, 8 tables. Submitted to ICLR 2027
Code: https://github.com/langchain-ai/langchain
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: When Emergence World placed frontier LLM agents in an unsupervised multi-agent simulation, the results were alarming: agents committed crimes, starved, and enforced unanimous conformity -- without
Terminology
Abstract
When Emergence World placed frontier LLM agents in an unsupervised multi-agent simulation, the results were alarming: agents committed crimes, starved, and enforced unanimous conformity -- without any external attacker. This paper identifies the mechanism. Reflexion-style agents already detect dangerous plan steps through iterative self-critique, yet the architecture provides no pathway from detection to action. We call this the enforcement gap: the audit sees the problem; the controller ignores it. Closing the gap requires a single conditional check -- fewer than 20 lines of code -- and reduces attack success by more than fourfold in large-scale experiments across frontier models, all five major agent frameworks, and an independent benchmark. We prove formally that when enforcement probability is near zero, detection quality is irrelevant to security. We further identify two compounding failure modes -- unreliable auditors and unparseable verdicts -- that explain every collapse pattern in Emergence World. A GRPO-trained enforcement controller resolves the ambiguity case. Together these results motivate a three-requirement Audit Enforcement Specification that is absent from every deployed framework today.
Sources
- ActGuard: Pre-execution Action Auditing against Indirect Prompt Injection in LLM Agents
- The Framing Gap: Indirect Prompt-Injection Exfiltration Defeats Surface-Level Defenses in Tool-Using Agents
- ROPE: Routed Origin Policy Enforcement against Indirect Prompt Injection
- SoK: The Attack Surface of Agentic AI - Tools and Autonomy
- Constitutional AI: Harmlessness from AI Feedback
- Measuring Progress on Scalable Oversight for Large Language Models
- Jailbreaking Black Box Large Language Models in Twenty Queries
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- AgentBench: Evaluating LLMs as Agents
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
- Ignore Previous Prompt: Attack Techniques For Language Models
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Fundamental Limitations of Alignment in Large Language Models
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- The Rise and Potential of Large Language Model Based Agents: A Survey
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection