Between the Commits: Process, Error, and Claim Reliability in a Wholly AI-Authored Codebase
cs.SE, cs.AI
Submitted: 2026-09-24
Updated: 2026-09-24
Code: https://github.com/doug-leith/extractor-code-development
Terminology
Sources
- Evaluating Large Language Models Trained on Code
- SWE-chat: Coding Agent Interactions From Real Users in the Wild
- Failure as a Process: An Anatomy of CLI Coding Agent Trajectories
- TrajAudit: Automated Failure Diagnosis for Agentic Coding Systems
- What Resolve Rate Hides: Trajectory Structure Diagnostics for Coding Agents
- A Large-Scale Comprehensive Measurement of AI-Generated Code in Real-World Repositories
- Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild
- AtomicCommitBench: Can Coding Agents Reconstruct Commit Histories from Squashed Patches?
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties