How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions
cs.SE, cs.AI, cs.HC
Submitted: 2026-05-28
Updated: 2026-08-31
Terminology
Sources
- Why Are AI Agent Involved Pull Requests (Fix-Related) Remain Unmerged? An Empirical Study
- SWE-chat: Coding Agent Interactions From Real Users in the Wild
- Composer 2 Technical Report
- Where Do AI Coding Agents Fail? An Empirical Study of Failed Agentic Pull Requests in GitHub
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- The Rise of AI Teammates in Software Engineering (SE) 3.0: How Autonomous Coding Agents Are Reshaping Software Engineering
- Understanding Code Agent Behaviour: An Empirical Study of Success and Failure Trajectories
- Beyond Resolution Rates: Behavioral Drivers of Coding Agent Success and Failure
- Position: Towards Bidirectional Human-AI Alignment
- Programming by Chat: A Large-Scale Behavioral Analysis of 11,579 Real-World AI-Assisted IDE Sessions
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties