Rouxii: Exploiting Honeypots with Deception-Aware AI Pentesters
cs.CR
Submitted: 2026-09-22
Updated: 2026-09-22
Code: https://github.com/cowrie/cowrie
License: http://creativecommons.org/licenses/by/4.0/
The gist: Honeypots are designed to deceive attackers, and recent work shows they can also derail autonomous LLM-based pentesters.
Terminology
Abstract
Honeypots are designed to deceive attackers, and recent work shows they can also derail autonomous LLM-based pentesters. These evaluations, however, largely consider attackers unaware of the deception they face. We study the opposite setting: an autonomous attacker explicitly equipped to recognize and act on honeypot fingerprints. We introduce Rouxii, an AI-driven penetration-testing framework that integrates counter-deception into reconnaissance and pivots from honeypot detection to exploitation. We evaluate matched vanilla and anti-deception Rouxii configurations across three reasoning models and eleven network setups over twelve cycles (1,544 attack reports). Between the matched cohorts, which differ only in the prompt, counter-deception raises correct honeypot identification from 19% to 97%, an effect strongest on OT services (11% to 97%), while false alarms on the real service stay at 0.7%. Deception-unaware baselines (PentestGPT, HackingBuddy) fail similarly, indicating the effect is not specific to our framework. Detection, moreover, is not the endpoint: through a white-box analysis of the honeypots themselves we show that a detected trap can be turned against its operator, demonstrating a denial-of-service that disables Conpot without tripping its liveness monitoring, and a corruption of the intelligence a GasPot instance reports. These findings show that deception effectiveness depends strongly on attacker knowledge, and that evaluations of honeypot resilience against AI attackers must account for adversaries that actively reason about and exploit the deception layer.
Sources
- EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security Vulnerabilities
- AutoPenBench: Benchmarking Generative Agents for Penetration Testing
- Hacking, The Lazy Way: LLM Augmented Pentesting
- What Makes a Good LLM Agent for Real-world Penetration Testing?
- LLM Agents can Autonomously Hack Websites
- Better Zero-Shot Reasoning with Role-Play Prompting
- The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
- An Empirical Evaluation of LLMs for Solving Offensive Security Challenges
- PentestAgent: Incorporating LLM Agents to Automated Penetration Testing
- Hacking Back the AI-Hacker: Prompt Injection as a Defense Against LLM-driven Cyberattacks
- From Sands to Mansions: Towards Automated Cyberattack Emulation with Classical Planning and Large Language Models
- Time-to-Lie: Identifying Industrial Control System Honeypots Using the Internet Control Message Protocol
- AutoAttacker: A Large Language Model Guided System to Implement Automatic Cyber-attacks
- ReAct: Synergizing Reasoning and Acting in Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs