Adversarial Testing of Automated Program Repair Agents for Security Vulnerabilities
cs.CR, cs.SE
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: 12 pages, 8 figures
Code: https://github.com/Aider-AI/aider
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- CodeR: Issue Resolving with Multi-Agent and Task Graphs
- Dissecting the SWE-Bench Leaderboards: Profiling Submitters and Architectures of LLM- and Agent-Based Repair Systems
- OpenHands: An Open Platform for AI Software Developers as Generalist Agents
- Red Teaming Program Repair Agents: When Correct Patches can Hide Vulnerabilities
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Keep the Conversation Going: Fixing 162 out of 337 bugs for $0.42 each using ChatGPT
- CodePlan: Repository-level Coding using LLMs and Planning
- QLPro: Automated Code Vulnerability Discovery via LLM and Static Code Analysis Integration
- VulEval: Towards Repository-Level Evaluation of Software Vulnerability Detection
- SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity
- SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code
- HardSecBench: Benchmarking the Security Awareness of LLMs for Hardware Code Generation
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs