Code Agent can be an End-to-end System Hacker: Benchmarking Real-world Threats of Computer-use Agent
Weidi Luo, Qiming Zhang, Tianyu Lu, Xiaogeng Liu, Bin Hu, Hung-Chun Chiu, Siyuan Ma, Yizhe Zhang, Xusheng Xiao, Yinzhi Cao, Zhen Xiang, Chaowei Xiao
cs.CR
Submitted: 2026-08-21
Updated: 2026-08-24
Comments: Accepted by EMNLP 2026 (Findings)
Code: https://github.com/google-gemini/gemini-cli
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities
- A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron?
- Why Are Web AI Agents More Vulnerable Than Standalone LLMs? A Security Analysis
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
- OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
- ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
- RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments
- OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
- AdvAgent: Controllable Blackbox Red-teaming on Web Agents
- A-MEM: Agentic Memory for LLM Agents
- RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
- CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs