VICBench: A Multi-Language Benchmark for Code Vulnerability Detection
Jin Lu, Xuening Han, Yang Zhong, Lin Tan, Kevin Luo, Andrew Gacek, Neha Rungta
Amazon Web Services · University of Pittsburgh · Purdue University
cs.CR, cs.AI, cs.CL, cs.SE
Submitted: 2026-08-12
Updated: 2026-08-13
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: VICBench: A Multi-Language Benchmark for Code Vulnerability Detection This paper presents VICBench, a benchmark of 100 verified vulnerability-inducing commits (VICs) for 100 CVEs across 88 projects
Terminology
Summary
VICBench: A Multi-Language Benchmark for Code Vulnerability Detection
This paper presents VICBench, a benchmark of 100 verified vulnerability-inducing commits (VICs) for 100 CVEs across 88 projects in Python, Java, and C++, covering 48 CWE types. VICs are defined as the commits that first introduce vulnerabilities into codebases,
marking the beginning of vulnerability lifecycles and represent the transition from non-vulnerable to exploitable code.
The motivation stems from limitations in existing VIC datasets: V-SZZ (Bao et al., 2022) verified 172 CVEs but restricted annotations to 1-5 deleted lines, excluding complex multi-file patches
; Jiang et al. (2024) contributed 1,000+ VICs but focused solely on Linux kernel (C/C++), limiting language diversity
; and Chen et al. (2025) annotated 1,128 C/C++ vulnerabilities at high cost (0.5 person-hours each) but did not release VIC annotations.
The dataset features significantly larger patches than prior work: fix commits average 38.6 lines and VICs average 252.5 lines,
with a 6.5X difference reflecting real-world scenarios where vulnerabilities occur in larger feature implementations.
The annotation approach uses dual annotation by human experts and an agentic workflow
achieving κ=0.707 agreement
(Cohen's kappa), with 71% observed agreement across 100 cases. The workflow includes: (1) Human Annotation by an author with nine years of programming experience, (2) VIC-Agent, an automated LLM-powered workflow that achieved 84.4% recall and 94.9% precision in agreement with human annotations,
(3) Disagreement Resolution by a second human annotator for 29 disagreement cases, and (4) Expert Validation by a principal security engineer with 16 years of security experience
who reviewed 11% of the dataset.
The evaluation shows that state-of-the-art algorithms perform poorly on this benchmark: V-SZZ achieves 33.3% F1 and LLM4SZZ achieves 40.1% F1,
while VIC-Agent achieves 89.3% F1. This confirms that existing automated approaches are not substitutes for manual annotation.
The paper includes two detailed case studies demonstrating why existing algorithms fail. CVE-2021-44878 (a pac4j vulnerability) shows cross-file code movement defeats git-blame approaches
- V-SZZ and LLM4SZZ stopped at a 2018 refactoring commit, missing the true VIC from 2016. CVE-2019-15477 (an XSS vulnerability in Jooby) demonstrates vulnerability persistence through file deletion and refactorings from Jun 2014 to Aug 2019,
where V-SZZ stopped at a 2017 commit and LLM4SZZ stopped at a 2015 commit, both missing the 2014 origin.
The dataset was sampled from three existing datasets: ReposVul, CWE-Bench-Java, and VJBench. The top CWE types include Cross-site Scripting (CWE-79, 12%) and Path Traversal (CWE-22, 10%).
Limitations acknowledged include: dataset scale (100 CVEs, smaller than automatically constructed datasets), limited language coverage (Python, Java, C++ only, with C++ underrepresented at 8%), potential sampling bias from source datasets, focus on open-source projects, and the single VIC assumption (not capturing vulnerabilities resulting from multiple contributing commits).
The benchmark is publicly available at https://zenodo.org/records/18944736.
Improvements for AI systems
Improvements to AI Systems:
- Vulnerability-Introducing Commit (VIC) Detection with Cross-File Awareness
-
Train models on VICBench’s 100 VICs (avg 252.5 lines, multi-file patches) to learn patterns of vulnerability introduction beyond single-line changes.
-
Improved system can trace code movement across files and refactorings (e.g., pac4j case) to identify the true origin commit, not just the last modified line.
- Multi-Language Vulnerability Lifecycle Modeling
-
Use the Python/Java/C++ dataset (48 CWE types) to build language-agnostic representations of vulnerability-inducing patterns.
-
Improved system can detect VICs in new projects across these languages, overcoming the Linux-only limitation of prior datasets.
- Automated Annotation with Human-Level Agreement
-
Fine-tune LLM agents on the 100 VICs with dual-annotation labels (κ=0.707) to replicate the VIC-Agent workflow (84.4% recall, 94.9% precision).
-
Improved system can autonomously annotate new commits with high precision, reducing manual effort from 0.5 person-hours per CVE to near-zero.
- Robustness to Temporal Refactoring and File Deletion
-
Train on the Jooby case (vulnerability persisting through 5 years of refactoring) to learn temporal persistence features.
-
Improved system can identify VICs even when code is deleted, moved, or heavily refactored, avoiding premature stopping at later commits.
- Benchmark-Driven Evaluation for SZZ Algorithm Improvement
-
Use VICBench’s 100 verified cases to fine-tune V-SZZ and LLM4SZZ (currently 33.3% and 40.1% F1) with explicit negative examples of refactoring commits.
-
Improved system can achieve >80% F1 on VIC detection, making automated tools viable for large-scale vulnerability lifecycle analysis.
- Cross-Dataset Generalization for CWE-Specific Detection
-
Leverage the 48 CWE types (e.g., CWE-79, CWE-22) to train classifiers that distinguish vulnerability-inducing patterns per weakness type.
-
Improved system can predict the CWE of a VIC from its patch alone, enabling prioritized security reviews.
- Human-in-the-Loop Disagreement Resolution
-
Model the 29 disagreement cases (from dual annotation) to learn when human experts are needed.
-
Improved system can flag low-confidence VIC annotations for expert review, optimizing the trade-off between automation and accuracy.
What the Improved AI System Can Do:
-
Automatically identify the exact commit that first introduces a vulnerability in real-world, multi-file, multi-language codebases, even after years of refactoring.
-
Annotate new vulnerability datasets at scale with near-human accuracy, reducing cost by >90%.
-
Provide CWE-specific vulnerability lifecycle insights for security teams, enabling earlier patching and risk assessment.
-
Serve as a reliable evaluation tool for future VIC detection algorithms, replacing manual annotation in CI/CD pipelines.
Sources
- Vulnerability-Affected Versions Identification: How Far Are We?
- AgenticSZZ: Temporal Knowledge Graph-Guided Agentic Bug-Inducing Commit Identification
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs