On the Reliability of LLM-Based Vulnerability Patching Benchmarks
cs.CR, cs.SE
Submitted: 2026-10-07
Updated: 2026-10-07
Code: https://github.com/ulikunitz/xz
Terminology
Sources
- Concerned with Data Contamination? Assessing Countermeasures in Code Language Model
- Large Language Models for Software Engineering: Survey and Open Problems
- Large Language Models for Software Engineering: A Systematic Literature Review
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
- SoK: Towards Effective Automated Vulnerability Repair
- VADER: A Human-Evaluated Benchmark for Vulnerability Assessment, Detection, Explanation, and Remediation
- Agent Lightning: Train ANY AI Agents with Reinforcement Learning
- What's in a Benchmark? The Case of SWE-Bench in Automated Program Repair
- ARVO: Atlas of Reproducible Vulnerabilities for Open-Source Software
- Coding Agents with Multimodal Browsing are Generalist Problem Solvers
- Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study
- Agentless: Demystifying LLM-based Software Engineering Agents
- Benchmark Data Contamination of Large Language Models: A Survey
- A Survey of LLM-based Automated Program Repair: Taxonomies, Design Paradigms, and Applications
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving
- AutoCodeRover: Autonomous Program Improvement
- Fixing Security Vulnerabilities with AI in OSS-Fuzz
- Learning to Retrieve from Agent Trajectories
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs