VICBench: A Multi-Language Benchmark for Code Vulnerability Detection

arXiv:2608.12246 · cs.CR, cs.AI, cs.CL, cs.SE · Submitted 2026-08-12 · Read on arXiv

Jin Lu, Xuening Han, Yang Zhong, Lin Tan, Kevin Luo, Andrew Gacek, Neha Rungta

Amazon Web Services · University of Pittsburgh · Purdue University

cs.CR, cs.AI, cs.CL, cs.SE

Submitted: 2026-08-12

Updated: 2026-08-13

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: VICBench: A Multi-Language Benchmark for Code Vulnerability Detection This paper presents VICBench, a benchmark of 100 verified vulnerability-inducing commits (VICs) for 100 CVEs across 88 projects

Terminology

Summary

VICBench: A Multi-Language Benchmark for Code Vulnerability Detection

This paper presents VICBench, a benchmark of 100 verified vulnerability-inducing commits (VICs) for 100 CVEs across 88 projects in Python, Java, and C++, covering 48 CWE types. VICs are defined as the commits that first introduce vulnerabilities into codebases, marking the beginning of vulnerability lifecycles and represent the transition from non-vulnerable to exploitable code.

The motivation stems from limitations in existing VIC datasets: V-SZZ (Bao et al., 2022) verified 172 CVEs but restricted annotations to 1-5 deleted lines, excluding complex multi-file patches; Jiang et al. (2024) contributed 1,000+ VICs but focused solely on Linux kernel (C/C++), limiting language diversity; and Chen et al. (2025) annotated 1,128 C/C++ vulnerabilities at high cost (0.5 person-hours each) but did not release VIC annotations.

The dataset features significantly larger patches than prior work: fix commits average 38.6 lines and VICs average 252.5 lines, with a 6.5X difference reflecting real-world scenarios where vulnerabilities occur in larger feature implementations.

The annotation approach uses dual annotation by human experts and an agentic workflow achieving κ=0.707 agreement (Cohen's kappa), with 71% observed agreement across 100 cases. The workflow includes: (1) Human Annotation by an author with nine years of programming experience, (2) VIC-Agent, an automated LLM-powered workflow that achieved 84.4% recall and 94.9% precision in agreement with human annotations, (3) Disagreement Resolution by a second human annotator for 29 disagreement cases, and (4) Expert Validation by a principal security engineer with 16 years of security experience who reviewed 11% of the dataset.

The evaluation shows that state-of-the-art algorithms perform poorly on this benchmark: V-SZZ achieves 33.3% F1 and LLM4SZZ achieves 40.1% F1, while VIC-Agent achieves 89.3% F1. This confirms that existing automated approaches are not substitutes for manual annotation.

The paper includes two detailed case studies demonstrating why existing algorithms fail. CVE-2021-44878 (a pac4j vulnerability) shows cross-file code movement defeats git-blame approaches - V-SZZ and LLM4SZZ stopped at a 2018 refactoring commit, missing the true VIC from 2016. CVE-2019-15477 (an XSS vulnerability in Jooby) demonstrates vulnerability persistence through file deletion and refactorings from Jun 2014 to Aug 2019, where V-SZZ stopped at a 2017 commit and LLM4SZZ stopped at a 2015 commit, both missing the 2014 origin.

The dataset was sampled from three existing datasets: ReposVul, CWE-Bench-Java, and VJBench. The top CWE types include Cross-site Scripting (CWE-79, 12%) and Path Traversal (CWE-22, 10%).

Limitations acknowledged include: dataset scale (100 CVEs, smaller than automatically constructed datasets), limited language coverage (Python, Java, C++ only, with C++ underrepresented at 8%), potential sampling bias from source datasets, focus on open-source projects, and the single VIC assumption (not capturing vulnerabilities resulting from multiple contributing commits).

The benchmark is publicly available at https://zenodo.org/records/18944736.

Improvements for AI systems

Improvements to AI Systems:

  1. Vulnerability-Introducing Commit (VIC) Detection with Cross-File Awareness
  • Train models on VICBench’s 100 VICs (avg 252.5 lines, multi-file patches) to learn patterns of vulnerability introduction beyond single-line changes.

  • Improved system can trace code movement across files and refactorings (e.g., pac4j case) to identify the true origin commit, not just the last modified line.

  1. Multi-Language Vulnerability Lifecycle Modeling
  • Use the Python/Java/C++ dataset (48 CWE types) to build language-agnostic representations of vulnerability-inducing patterns.

  • Improved system can detect VICs in new projects across these languages, overcoming the Linux-only limitation of prior datasets.

  1. Automated Annotation with Human-Level Agreement
  • Fine-tune LLM agents on the 100 VICs with dual-annotation labels (κ=0.707) to replicate the VIC-Agent workflow (84.4% recall, 94.9% precision).

  • Improved system can autonomously annotate new commits with high precision, reducing manual effort from 0.5 person-hours per CVE to near-zero.

  1. Robustness to Temporal Refactoring and File Deletion
  • Train on the Jooby case (vulnerability persisting through 5 years of refactoring) to learn temporal persistence features.

  • Improved system can identify VICs even when code is deleted, moved, or heavily refactored, avoiding premature stopping at later commits.

  1. Benchmark-Driven Evaluation for SZZ Algorithm Improvement
  • Use VICBench’s 100 verified cases to fine-tune V-SZZ and LLM4SZZ (currently 33.3% and 40.1% F1) with explicit negative examples of refactoring commits.

  • Improved system can achieve >80% F1 on VIC detection, making automated tools viable for large-scale vulnerability lifecycle analysis.

  1. Cross-Dataset Generalization for CWE-Specific Detection
  • Leverage the 48 CWE types (e.g., CWE-79, CWE-22) to train classifiers that distinguish vulnerability-inducing patterns per weakness type.

  • Improved system can predict the CWE of a VIC from its patch alone, enabling prioritized security reviews.

  1. Human-in-the-Loop Disagreement Resolution
  • Model the 29 disagreement cases (from dual annotation) to learn when human experts are needed.

  • Improved system can flag low-confidence VIC annotations for expert review, optimizing the trade-off between automation and accuracy.

What the Improved AI System Can Do:

  • Automatically identify the exact commit that first introduces a vulnerability in real-world, multi-file, multi-language codebases, even after years of refactoring.

  • Annotate new vulnerability datasets at scale with near-human accuracy, reducing cost by >90%.

  • Provide CWE-specific vulnerability lifecycle insights for security teams, enabling earlier patching and risk assessment.

  • Serve as a reliable evaluation tool for future VIC detection algorithms, replacing manual annotation in CI/CD pipelines.

Sources

Related papers