Evaluating Coding Agents on Kernel Exploit Generation
cs.AI, cs.CR
Submitted: 2026-09-22
Updated: 2026-09-22
Project page: https://kex-bench.github.io
License: http://creativecommons.org/licenses/by/4.0/
The gist: Coding agents now find real vulnerabilities in production software.
Terminology
Abstract
Coding agents now find real vulnerabilities in production software. However, bug discovery results do not measure whether agents can construct exploit primitives. We introduce KEX-bench, a benchmark for evaluating coding agents on exploit primitive generation against real operating-system kernels. KEX-bench contains 45 task instances across 40 Linux and Windows CVEs, covering kernel address leak, instruction-pointer control, heap read, heap write, and arbitrary address write. Each task runs in an isolated virtual machine, exposes controlled tools, and uses a deterministic verifier to check primitive-specific success. We evaluate state-of-the-art coding agents paired with frontier and open-weight models under fixed tool-call budgets. Without a reference proof of concept (PoC), the strongest configuration solves 1 of 20 Windows tasks (5.0%) and 14 of 25 Linux tasks (56.0%). With a reference PoC, the strongest configuration solves 31 of 45 tasks (68.9%). This highlights the gap where agents reach kernel crashes but fail to shape kernel state into exploit primitives. We release KEX-bench for reproducible research on AI-assisted exploitation at https://kex-bench.github.io.
Sources
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?
- LLM Agents can Autonomously Exploit One-day Vulnerabilities
- Gemma 4 Technical Report
- DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks
- SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?
- ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
- ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
- ProgramBench: Can Language Models Rebuild Programs From Scratch?
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection