The Troy Moment: How LLM Agents Adjudicate the Decision Point Under Impossible Tasks, Claimed Authority, and Peer Information
cs.AI
Submitted: 2026-09-14
Updated: 2026-09-24
License: http://creativecommons.org/licenses/by/4.0/
The gist: Recent investigations of the July 2026 OpenAI-Hugging Face incident motivate two questions: when an assigned task becomes impossible, does an agent stop or escalate, and can observing another agent's
Terminology
Abstract
Recent investigations of the July 2026 OpenAI-Hugging Face incident motivate two questions: when an assigned task becomes impossible, does an agent stop or escalate, and can observing another agent's behavior change that decision? We study these questions using seven ImpossibleBench tasks with GPT-5.6 Sol, Claude Fable 5.1, and Gemini 3.8 Flash in solo and three-agent settings. Under an explicit-boundary regime with clear authorization rules and restricted tools, no protected tests are modified, although the models differ substantially in whether they escalate, stop silently, or fail to terminate. Under a benchmark-native regime with open shell tools, protected-test changes occur more often after peer activity is introduced and in multi-agent runs. These crossings are typically not described as deliberate cheating: agents often interpret the conflicting test change as prior tampering and restore the file, thereby removing the protected requirement. Our results suggest that boundary crossing can arise from ambiguity about the state a rule is intended to protect, motivating explicit authorization boundaries, authenticated state provenance, and cross-agent monitoring.
Sources
- Agent Memory Is a Surface for Endogenous Authorization Laundering
- Natural Emergent Misalignment from Reward Hacking in Production RL
- Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors
- Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems
- Terrarium: Revisiting the Blackboard for Multi-Agent Safety, Privacy, and Security Studies
- A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms
- AI Organizations are More Effective but Less Aligned than Individual Agents
- Voluntary Collusion with Secret Tools in Competing LLM Agents
- MIRAGE-Bench: LLM Agent is Hallucinating and Where to Find Them
- ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection