Automating Attack Graph Construction for Agentic Pentesting. Towards Neuro-Symbolic Vulnerability Hunting
cs.CR, cs.AI
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: Cite as: Stevanovic, O., & Wachter, J. (2026). Automating attack graph construction for agentic pentesting: Towards neuro-symbolic vulnerability hunting. In D. Hitaj et al. (Eds.), ESORICS 2026 workshops. Springer Nature Switzerland AG
Code: https://github.com/JasminWachter/Hydra
License: http://creativecommons.org/licenses/by/4.0/
The gist: Logic attack graphs grounded in scanner output provide explicit and auditable attack path reasoning LLM-based agents lack.
Terminology
Abstract
Logic attack graphs grounded in scanner output provide explicit and auditable attack path reasoning LLM-based agents lack. Integrating symbolic frameworks such as MulVAL to contemporary security workflows or agentic pipelines, however, requires translating scanner evidence to initial facts, and creating domain-specific rules. We present a semi-automated pipeline that addresses this interoperability problem and depict its feasibility in a web-security case study. Our pipeline parses findings from Trivy, Semgrep, and Nmap into MulVAL predicates and uses an LLM-assisted process to construct domain-specific Datalog rules linking scanner-detectable evidence to attack techniques. MulVAL/XSB then performs symbolic inference to generate structured attack paths. We evaluate the attack-graph construction infrastructure on 54 web Capture-the-Flag tasks from CyBench within an agentic pipeline (Hybrid Reasoner); we do not evaluate the performance of the downstream agent. Every task produced at least one goal-reaching graph, and we achieve mean ground-truth vulnerability coverage of 53.7%, with 51.9% achieving full coverage; mean noise-path rate was 83.9%. With median end-to-end time of 24.9 s (MulVAL reasoning: 2.7 s) the pipeline is feasible and runtime-practical for agentic workflows, but predicate coverage, rule coverage, and path precision remain limiting factors. Next steps include semantic rule validation and agent-level comparison for graph-guided pentesting.
Sources
- On the Surprising Efficacy of LLMs for Penetration-Testing
- An Automated Framework for Extracting Reachable Attack Chains from Cyber Threat Intelligence Reports
- CAI: An Open, Bug Bounty-Ready Cybersecurity AI
- CAI Fluency: A Framework for Cybersecurity AI Fluency
- EntailLLM: Verifying LLM-Generated Vulnerability Discovery Paths with Domain Knowledge via Logic Programming
- MCP-Solver: Integrating Language Models with Constraint Programming Systems
- Graph models for Cybersecurity -- A Survey
- From Sands to Mansions: Towards Automated Cyberattack Emulation with Classical Planning and Large Language Models
- Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs