PatchBench: Evaluating AI Agents for Vulnerability Patching
cs.CR, cs.AI, cs.SE
Submitted: 2026-09-03
Updated: 2026-09-03
Code: https://github.com/ai-sec-lab/PatchBench
Terminology
Sources
- MarsCode Agent: AI-native Automated Bug Fixing
- CodeBLEU: a Method for Automatic Evaluation of Code Synthesis
- PATCHEVAL: A New Benchmark for Evaluating LLMs on Patching Real-World Vulnerabilities
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs