AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
cs.AI
Submitted: 2026-04-03
Updated: 2026-09-22
Project page: https://yunhao-feng.github.io/AgentHazard
License: http://creativecommons.org/licenses/by/4.0/
The gist: Computer-use agents extend language models from text generation to persistent action over tools, files, and execution environments.
Terminology
Abstract
Computer-use agents extend language models from text generation to persistent action over tools, files, and execution environments. Unlike chat systems, they maintain state across interactions and translate intermediate outputs into concrete actions. This creates a distinct safety challenge in that harmful behavior may emerge through sequences of individually plausible steps, including intermediate actions that appear locally acceptable but collectively lead to unauthorized actions. We present AgentHazard, a benchmark for evaluating harmful behavior in computer-use agents. AgentHazard contains 2,653 instances spanning diverse risk categories and attack strategies. Each instance pairs a harmful objective with a sequence of operational steps that are locally legitimate but jointly induce unsafe behavior. The benchmark evaluates whether agents can recognize and interrupt harm arising from accumulated context, repeated tool use, intermediate actions, and dependencies across steps. We evaluate AgentHazard on Claude Code, OpenClaw, and IFlow using mostly open or openly deployable models from the Qwen3, Kimi, GLM, and DeepSeek families. Our experimental results indicate that current systems remain highly vulnerable. In particular, when powered by Qwen3-Coder, Claude Code exhibits an attack success rate of 73.63%, suggesting that model alignment alone does not reliably guarantee the safety of autonomous agents.
Sources
- Anthropic Economic Index report: Uneven geographic and enterprise AI adoption
- Qwen Technical Report
- LLM-Safety Evaluations Lack Robustness
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
- LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering
- Qwen3Guard Technical Report
- Benchmarking Correctness and Security in Multi-Turn Code Generation
- Don't Let the Claw Grip Your Hand: A Security Analysis and Defense Framework for OpenClaw
- Kimi K2: Open Agentic Intelligence
- Let It Flow: Agentic Crafting on Rock and Roll, Building the ROME Model within an Open Agentic Learning Ecosystem
- OpenHands: An Open Platform for AI Software Developers as Generalist Agents
- From Agent-Only Social Networks to Autonomous Scientific Research: Lessons from OpenClaw and Moltbook, and the Architecture of ClawdLab and Beach.Science
- Internal Safety Collapse in Frontier Large Language Models
- Qwen3 Technical Report
- GLM-5: from Vibe Coding to Agentic Engineering
- Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
- Agent-SafetyBench: Evaluating the Safety of LLM Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection