Honeyquest for LLMs: Rethinking Cyber Deception for AI Attackers
Kerri Prinos, Lilianne Brush, Cameron Denton
cs.CR, cs.CL
Submitted: 2026-08-17
Updated: 2026-08-18
Comments: 20 pages, 4 figures, 2 tables
Project page: https://exploitbench.ai
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: The empirical foundation of cyber deception relies on human-centered hypotheses, but the rapid emergence of autonomous, AI-enabled attackers challenges whether this foundation transfers to AI agents.
Terminology
Abstract
The empirical foundation of cyber deception relies on human-centered hypotheses, but the rapid emergence of autonomous, AI-enabled attackers challenges whether this foundation transfers to AI agents. To address this, we introduce an automated evaluation framework adapted from the Honeyquest instrument to assess LLM attacker judgment at scale. Our 21-LLM cohort spanned 10 providers, diverse architectures and specializations, open- and closed-weight models, and parameter scales from 8B to over 1T. We evaluated the performance of this LLM cohort (yielding 10,962 responses) against the 47-participant human baseline across an identical set of 174 reconnaissance queries. Our empirical evaluation reveals three key findings that establish LLMs as a distinct attacker class: (1) every model in our cohort falls for deceptive traps at a significantly higher rate than human attackers; (2) the defensive attention-diversion effect observed in humans is statistically absent in our LLM cohort; and (3) a critical recognition-action gap, where LLMs successfully articulate trap recognition in their reasoning but exploit the deceptive elements anyway 73.4% of the time. Across the 21 models, trap recognition in reasoning text did not predict fell-for-trap behavior (Spearman r = +0.08, p = 0.73). Ultimately, these findings demonstrate that human-centered deception hypotheses do not reliably transfer to AI attackers, highlighting the critical need for new research into AI-native active defense frameworks.
Sources
- Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies
- Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- LLM Agents can Autonomously Exploit One-day Vulnerabilities
- AI Agents Enable Adaptive Computer Worms
- Machine Psychology
- ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
- Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- gpt-oss-120b & gpt-oss-20b Model Card
- LLM Agent Honeypot: Monitoring AI Hacking Agents in the Wild
- Agents of Chaos
- Kimi K2.5: Visual Agentic Intelligence
- Qwen3 Technical Report
- Llama-3.1-FoundationAI-SecurityLLM-Reasoning-8B Technical Report
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs