Trust but Verify? Uncovering the Security Debt of Autonomous Coding Agents
A H M Nazmus Sakib, Dipayan Banik, Murtuza Jadliwala
cs.CR
Submitted: 2026-07-14
Comments: Accepted at the KDD 2026 Workshop on Agentic Software Engineering (AgenticSE)
Code: https://github.com/nazanaza2970/agentic_se_kdd_submission
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: The increasing adoption of autonomous coding agents accelerates software development but also introduces scoped security risks within high-impact file paths that can outpace traditional human review
Terminology
Abstract
The increasing adoption of autonomous coding agents accelerates software development but also introduces scoped security risks within high-impact file paths that can outpace traditional human review capacity. While prior research has primarily evaluated these systems in terms of functional correctness and productivity, this paper presents a large-scale empirical study using the AIDev dataset to systematically characterize security code smells in agent-generated pull requests (PRs). Through a combination of a validated LLM-as-a-judge framework and manual qualitative analysis, we identify and classify security misconfigurations across 16,112 file changes spanning 4,022 pull requests. Our results reveal that 38.9% of agent-generated PRs contain at least one security smell, with supply chain integrity issues accounting for 82.3% of all detected security smells. Furthermore, hard-coded credentials constitute 99.6% of all critical-severity security smells. Crucially, we find that human collaborators are responsible for introducing 67.6% of genuine leaked secrets within these agent-assisted workflows, while existing automated and human review processes fail to detect 81.1% of these credentials prior to integration. These findings highlight substantial security risks in agent-assisted software development workflows and suggest a potential reduction in developer vigilance. They also underscore the urgent need for context-aware security guardrails implemented directly at the point of human-AI collaboration.
Sources
- Evaluating Large Language Models Trained on Code
- Understanding Dominant Themes in Reviewing Agentic AI-authored Code
- Do Autonomous Agents Contribute Test Code? A Study of Tests in Agentic Pull Requests
- Artificial Intelligence in Open Source Software Engineering: A Foundation for Sustainability
- The Rise of AI Teammates in Software Engineering (SE) 3.0: How Autonomous Coding Agents Are Reshaping Software Engineering
- How AI Coding Agents Modify Code: A Large-Scale Study of GitHub Pull Requests
- Good Vibrations? A Qualitative Study of Co-Creation, Communication, Flow, and Trust in Vibe Coding
- A Task-Level Evaluation of AI Agents in Open-Source Projects
- Agentic AI Software Engineers: Programming with Trust
- Assessing the Quality and Security of AI-Generated Code: A Quantitative Analysis
- On the Use of Agentic Coding: An Empirical Study of Pull Requests on GitHub
- Testing with AI Agents: An Empirical Study of Test Generation Frequency, Quality, and Coverage
- Let's Make Every Pull Request Meaningful: An Empirical Analysis of Developer and Agentic Pull Requests
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs