AgenticRepair: Multi-Faceted Program Context Engineering for Agentic Vulnerability Repair
Michael Fu, Qiyue Mei, Patanamon Thongtanunam, Kla Tantithamthavorn
cs.SE, cs.AI, cs.CR
Submitted: 2026-07-31
Comments: Under Review at IEEE TSE
Code: https://github.com/mity/md4c
Project page: https://sec-bench.github.io
License: http://creativecommons.org/licenses/by/4.0/
The gist: Automated vulnerability repair aims to reduce the time and effort required to patch security flaws from a vulnerability triage report.
Terminology
Abstract
Automated vulnerability repair aims to reduce the time and effort required to patch security flaws from a vulnerability triage report. Recent agentic AI approaches have shown promising results in automated program repair. However, vulnerability repair demands richer program context than general bug repair - context that security engineers routinely assemble in practice but that existing agentic approaches do not engineer. We identify three critical gaps: code-structure context capturing cross-file data flows and memory operation patterns, runtime-execution context revealing crash semantics and memory origins, and commit-history context recovering how fragile code patterns were introduced. We present AgenticRepair, an agentic vulnerability repair framework that addresses the gaps through multi-faceted program context engineering. AgenticRepair orchestrates three specialized LLM subagents to engineer the contexts, which are then embedded into the memory of a dedicated repair subagent for context-conditioned patch synthesis. Evaluated on SEC-Bench comprising 300 real-world instances with sanitizer-based patch verification, AgenticRepair achieves a 73% success rate, substantially outperforming the strongest baseline by 29%. Our ablation study confirms that the three context facets are mutually complementary, and that multi-agent scaffolding and base-model capacity each play an essential role. Collectively, these findings establish multi-faceted program context engineering as a promising design direction for agentic vulnerability repair.
Sources
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
- DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems
- ARVO: Atlas of Reproducible Vulnerabilities for Open-Source Software
- OpenHands: An Open Platform for AI Software Developers as Generalist Agents
- Agentless: Demystifying LLM-based Software Engineering Agents
- Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties