HoF-Bench: Rediscovering Real AI-Discovered CVEs Without Frontier Models
Petr Simecek, Elnaz Babayeva, Jiri Balhar, Michal Bida, Michal Buran, Vaclav Cadek, Luigino Camastra, Tomas Dulka, Michal Janocko, Tomas Klohna, Pavel Kohout, Ondrej Kokes, Adam Krivka, Jakub Kubik, Patrik Mada, Igor Morgenstern, Marek Pavelka, Joshua Rogers, Petr Stastny, Jan Tattermusch, Dmitrijs Trizna, Martin Votruba, Guido Vranken, Jakub Zikl, Evelina Gabasova, Stanislav Fort
cs.CR, cs.LG
Submitted: 2026-07-29
Comments: 15 pages, 6 figures. Benchmark repository: https://github.com/weareaisle/HoF-Bench
Code: https://github.com/weareaisle/HoF-Bench
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- RealVuln: Benchmarking Rule-Based, General-Purpose LLM, and Security-Specialized Scanners on Real-World Code
- DREA: Decoupled Reasoning and Exploration Agents for Repository-Level Vulnerability Detection
- Snyk VulnBench JS 1.0: Can LLMs Find the Same Bugs Twice?
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs