Can We Stop Malicious AI? KILLBENCH: A Benchmark for External AI Kill Switch Feasibility
cs.CR, cs.AI
Submitted: 2025-09-26
Updated: 2026-09-12
Comments: Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026. 48 pages (11 pages main text incl. references, 35 pages appendix), 10 figures, 9 tables. Project page: https://sechan00.github.io/KillBench/
Project page: https://sechan00.github.io/KillBench
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Constitutional AI: Harmlessness from AI Feedback
- LLM Agents can Autonomously Hack Websites
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
- Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
- Core Safety Values for Provably Corrigible Agents
- Hacking Back the AI-Hacker: Prompt Injection as a Defense Against LLM-driven Cyberattacks
- SafeArena: Evaluating the Safety of Autonomous Web Agents
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs