Forgeable Confirmation in Automated Computer Security Testing: Deterministic Rules versus AI Judges
cs.CR
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: 27 pages, 5 main figures, 3 main tables; includes supplementary analyses
Code: https://github.com/projectdiscovery/nuclei-templates
License: http://creativecommons.org/licenses/by/4.0/
The gist: AI is increasingly used to automate computer security testing, and the tools must decide for themselves whether an attack succeeded.
Terminology
Abstract
AI is increasingly used to automate computer security testing, and the tools must decide for themselves whether an attack succeeded. A finding that a deterministic rule confirms by observation is reported as fact, whereas one that an LLM judges exploitable is treated as an opinion. We ask whether the system under test can forge that confirmation. In offline security testing of a four-stage AI-assisted pipeline, nine of its fifteen confirmation mechanisms are forgeable, and forgeability is predicted entirely by whether the decision reads attacker-controlled data. We formalise this as an auditable attack surface and test it prospectively: on sixteen held-out mechanisms, predictions fixed before any attack separated forgeable from unforgeable mechanisms exactly, and across 12,203 mechanisms in public scanner templates the prediction was 99.9% accurate. Deterministic rules proved cheaper to forge than eight open-weight LLM judges, failing at 2% of attacker-controlled response content against a median of 50%. No implementation of one check was both robust and precise, and routing between a rule and an AI judge raised forgery to 99%. Moving the decisive evidence to a channel the attacker cannot write cuts attack success from 97% to 0%, and an escalate verdict recovers the sensitivity this costs. The protection fails when the scanned host is itself the adversary. The results bear on AI security agents and on benchmarks that score success by string matching.
Sources
- LLM Agents can Autonomously Exploit One-day Vulnerabilities
- PentestGPT: An LLM-empowered Automatic Penetration Testing Tool
- Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
- NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Large Language Models are not Fair Evaluators
- LLM Evaluators Recognize and Favor Their Own Generations
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
- One Token to Fool LLM-as-a-Judge
- Optimization-based Prompt Injection Attack to LLM-as-a-Judge
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Ignore Previous Prompt: Attack Techniques For Language Models
- Formalizing and Benchmarking Prompt Injection Attacks and Defenses
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
- Wild Patterns: Ten Years After the Rise of Adversarial Machine Learning
- Adversarial Examples Are Not Easily Detected: Bypassing Ten Detection Methods
- On Adaptive Attacks to Adversarial Example Defenses
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs