QRS: A Rule-Synthesizing Neuro-Symbolic Triad for Autonomous Vulnerability Discovery
George Tsigkourakos, Constantinos Patsakis
cs.CR
Submitted: 2026-08-20
Updated: 2026-08-21
Code: https://github.com/aio-libs/aioh
Project page: https://hugovk.github.io/top-pypi-packages
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Static Application Security Testing (SAST) tools are integral to DevSecOps pipelines, yet tools like CodeQL, Semgrep, and SonarQube remain constrained: they require expert-crafted queries, generate
Terminology
Abstract
Static Application Security Testing (SAST) tools are integral to DevSecOps pipelines, yet tools like CodeQL, Semgrep, and SonarQube remain constrained: they require expert-crafted queries, generate excessive false positives, and detect only predefined patterns. Recent work augments SAST with Large Language Models (LLMs), but typically requires fine-tuning or uses LLMs only to triage outputs rather than reason directly about vulnerability semantics. We introduce QRS (Query, Review, Sanitize), a neuro-symbolic framework that inverts this paradigm by relocating the LLM to the pipeline's generative core. Rather than filtering static-rule results, QRS employs three autonomous agents that generate CodeQL queries from a structured schema and few-shot examples, then validate findings through semantic reasoning and minimal proof-of-concept exploits, manually verified. This lets QRS surface vulnerability classes beyond predefined patterns while reducing the code volume needing audit. We evaluate QRS on complete packages rather than isolated snippets, across Python and Java ecosystems and three datasets. On the 100 most-downloaded PyPI packages, QRS reaches up to 94.06% verdict accuracy and surfaces 41 medium-to-high vulnerabilities: 8 received new CVEs, 4 were acknowledged via documentation updates, and the remaining 29 were previously published CVEs rediscovered from source alone. To evaluate language portability and enable direct comparison with prior work, we extend QRS to Java and evaluate it on CWE-Bench-Java, detecting 58/90 in-scope CVEs (64.44%) at 87.60% accuracy, 98.79% recall, and 0.857 F1 (0.794 macro-averaged) in one scan. QRS achieves these results with low time overhead and manageable token costs while handling large codebases, showing that LLM-driven query synthesis and review can complement curated rule sets and surface patterns that evade existing industry tools.
Sources
- LLM-Driven SAST-Genius: A Hybrid Static Analysis Framework for Comprehensive and Actionable Security
- QLPro: Automated Code Vulnerability Discovery via LLM and Static Code Analysis Integration
- Neuro-symbolic Static Analysis with LLM-generated Vulnerability Patterns
- The Malware as a Service ecosystem
- LLM4Vuln: A Unified Evaluation Framework for Decoupling and Enhancing LLMs' Vulnerability Reasoning
- From CVE Entries to Verifiable Exploits: An Automated Multi-Agent Framework for Reproducing CVEs
- QLCoder: A Query Synthesizer For Static Analysis of Security Vulnerabilities
- VulAgent: Hypothesis-Validation based Multi-Agent Vulnerability Detection
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs