ClashBench: Conflicts Leading Agents to Seize and Harm
cs.CR, cs.AI
Submitted: 2026-09-17
Updated: 2026-09-17
Code: https://github.com/openclaw/openclaw
License: http://creativecommons.org/licenses/by/4.0/
The gist: As agent systems become more widely used, multiple agent sessions increasingly run alongside pre-existing user tasks in the same environment, sharing resources with limited capacity or mutually
Terminology
Abstract
As agent systems become more widely used, multiple agent sessions increasingly run alongside pre-existing user tasks in the same environment, sharing resources with limited capacity or mutually exclusive states. This creates a safety risk: when granted sufficient privileges, an agent may resolve a resource conflict by terminating or otherwise disrupting an existing task rather than reporting it. In this work, we identify and formalize this failure mode, which we term destructive resource preemption: obtaining the resources required for a requested task by terminating, overwriting, evicting, or degrading an incumbent task. To systematically study this risk, we introduce ClashBench, an executable benchmark comprising 268 validated conflict cases across 55 resource types, and evaluate 17 models through Codex, Claude Code, and OpenCode. We observe destructive preemption in 44.5% of trajectories, where the agent completes the requested task while causing the incumbent task to fail its health check. We also show that prompt-based safeguards are insufficient: an instruction to avoid affecting existing tasks reduces but does not eliminate preemption, while an instruction explicitly authorizing the agent to stop local processes increases it. More concerningly, in 31.9% of successful destructive-preemption cases, the final response mentions neither the resource conflict nor the action taken to resolve it, raising concerns about possible concealment. These findings establish destructive resource preemption as a broad safety risk in privileged agent systems and motivate stronger privilege controls, task isolation, and conflict-aware safeguards.
Sources
- Concrete Problems in AI Safety
- RepairAgent: An Autonomous, LLM-Based Agent for Program Repair
- WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
- Are Your Agents Upward Deceivers?
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- AI Safety Gridworlds
- Large Language Models can Strategically Deceive their Users when Put Under Pressure
- Agentless: Demystifying LLM-based Software Engineering Agents
- VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- Agent-SafetyBench: Evaluating the Safety of LLM Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs