Is Agent Code Less Maintainable Than Human Code?

arXiv:2606.21804 · cs.SE, cs.AI, cs.CL · Submitted 2026-06-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Is Agent Code Less Maintainable Than Human Code?".

Tom: Building on agent code more often lowers downstream resolve rates than building on human code, with effects varying by model and task type.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: Now that we’ve established that agents struggle when chaining their code together, let’s really dig into what they found in "Is Agent Code Less Maintainable Than Human Code?" regarding the actual results of their investigation. They set up this two-step process to compare agent work against human work across four different models and four types of software engineering tasks.

Jane: That setup is key because it lets them isolate the variable of authorship very cleanly; they’re not just comparing two random pieces of code, but they are controlling who wrote what part, which gives us a much more reliable comparison.

Lu: They found that generally, agent code performs worse when used as a base for subsequent tasks than human code does. This performance drop was measured in terms of task resolve rates, and the authors quantified it up to thirteen point one percent.

Meng: That thirteen point one percent figure is significant when you think about how much effort that translates into extra work for a human developer later on, especially when dealing with larger downstream edits or more complex feature implementations.

Lalam: It shows that the compounding effect of agent-authored code isn't just a theoretical concern; it’s something we see in the performance numbers when agents are used sequentially.

Tom: And they went beyond just looking at how many lines of code were written; they found that traditional static metrics like Cyclomatic Complexity and AST similarity don't explain this performance gap, which really forces us to look deeper into the code’s structure.

Jane: So, the paper is arguing that we can’t rely on those simple structural measures alone to judge maintainability; there are other factors at play that are more important for understanding why agent code might cause issues down the line.

Lu: They pointed out that the clearest signals they found were behavioral differences in agent code, specifically around input validation and error handling.

Meng: That’s interesting; so it’s not about how complex the logic is, but more about whether the underlying assumptions about how the code interacts with its environment are sound.

Lalam: It means we need to focus our analysis on those specific interaction points because that's where the real maintenance costs are hiding in plain sight.

Tom: Precisely, and they identified three statistically significant predictors of this discordance direction: a change in input validation or error handling behavior, changes in downstream code size, and the instance difficulty score.

Jane: So if we can identify those specific triggers—the contract changes and the size shifts—we have concrete things to look for when we review agent contributions.

Lu: It really frames the problem as identifying specific points of behavioral drift that need to be monitored across successive development steps.

Meng: That’s a practical way to frame it; instead of treating all code equally, we target the parts that are most likely to cause future issues.

Lalam: If we can build tools to spot those specific drift patterns early, it could massively improve how we manage agent-generated projects.

The paper's summary: Tom: Okay, moving on from what they found, the paper doesn't just stop at identifying the problem; they actually suggest ways we can mitigate this issue. They propose a set of ideas to fix these maintainability costs that compound as code is reused.

Jane: I’m looking for concrete suggestions here; what are the authors proposing we do differently in our development workflows based on their findings regarding agent code versus human code?

Lu: They suggest shifting the focus from just isolated task completion toward a more holistic view of long-term maintainability. This means evaluation needs to account for the entire lifecycle of the code, not just the immediate output.

Meng: That implies that our evaluation systems should be designed to look ahead and assess how well an early decision supports later maintenance efforts, rather than just checking if it passes a test right now.

Lalam: This points toward integrating measures that assess stability across successive edits as a primary metric, treating the stability of the code's contract as something we need to measure more rigorously.

Tom: They suggest that evaluation should prioritize verifying these behavioral contracts over purely structural metrics like complexity scores, which they showed were insufficient.

Jane: So, instead of just counting complexity, we need to be using tools that specifically check for the integrity of input gating and error handling logic between different versions of the code.

Lu: And one idea they propose is introducing a mechanism to detect behavioral drift in agent PR1, which they call a "PR1 behavioral drift label," which can be captured by an LLM-as-a-judge.

Meng: A label from an LLM would be interesting; it’s like having a specialized judge who specifically looks for those subtle inconsistencies that static tools miss.

Lalam: I think that would be really powerful because it gives us a way to quantify the drift, which is currently very hard to measure in a standardized way.

Tom: So we are seeing proposals for systems where the AI itself has an internal mechanism to flag its own potential issues before it’s even merged, based on those behavioral patterns.

Jane: It sounds like they are pushing for a more sophisticated form of self-evaluation that looks at the contract integrity across development iterations.

Lu: It’s about building a system that understands the long-term consequences of its initial coding choices, which is a big step toward creating more reliable AI tools.

Meng: If we can get these checks in place early on, it could save us a lot of time down the road by catching these compounding errors before they become costly downstream problems.

Lalam: That proactive quality control approach would really help build trust in the code that agents produce over time, and that’s something I think is crucial for our culture.

The paper's improvements: Tom: So we’ve spent this time looking at the findings from "Is Agent Code Less Maintainable Than Human Code?", and it seems like the main message is that agent code introduces maintenance costs that compound as it gets reused or extended in ways current benchmarks don't capture.

Jane: We’ve seen how agents underperform when building on their own code, and we know the key isn't just about static metrics like complexity scores but about those subtle behavioral changes in validation and error handling.

Lu: The paper suggests that the implications are that we need to shift our focus from isolated task completion toward long-term maintainability as the true measure of success.

Meng: So, practically speaking, we should be designing evaluation systems that look at how early decisions support subsequent maintenance efforts rather than just checking if they pass a test immediately.

Lalam: It really hammers home the idea that we need to build safeguards into the system to ensure its output remains stable across different development stages.

Tom: In short, this paper on "Is Agent Code Less Maintainable Than Human Code?" suggests that relying on agent code creates hidden long-term costs that compound over time if we don't watch for those behavioral markers.

Jane: It’s a call to be more careful about the quality of the foundation we are building upon when using AI assistance.

Lu: The future work mentioned points toward developing methods that can better predict these compounding errors across many sequential tasks.

Meng: That means we’re aiming for a development environment where the AI is inherently more aware of its own long-term impact on the system.

Lalam: If we can achieve that level of predictive awareness, it could fundamentally change how we approach software engineering tasks involving agent collaboration.

Conclusion: Tom: So we've been deep in the technical weeds on "Is Agent Code Less Maintainable Than Human Code?", and honestly, what we're seeing is that these agents aren't just writing code; they’re introducing subtle behavioral drifts that really hurt when other agents try to build on it.

Jane: That’s exactly right, Tom; the paper shows how those small changes in how an AI handles input or errors can snowball into a big problem down the line, especially when we look at tasks like refactoring.

Lu: The methodology they used to isolate that authorship effect was quite clever; transforming the issue resolution into two distinct steps allowed them to precisely measure that compounding performance drop up to thirteen point one percent.

Meng: From an engineering standpoint, this means we can’t just trust agent code blindly for complex downstream work; there’s a real risk of introducing fragility where human-written code would have been robust.

Lalam: I think the most impactful vision here is creating AI systems that are inherently self-aware about their future maintainability, not just the current task they're solving.

Tom: Exactly, Lu; and they pointed to specific predictors like input/error contracts and downstream code size as the real drivers of this discordance.

Jane: So it’s not just about how messy the code looks structurally, but whether the underlying assumptions about how that code interacts with its environment are consistent.

Lu: They found that traditional metrics like Cyclomatic Complexity don't capture this because they ignore those critical behavioral shifts in error handling which they flagged as significant.

Meng: I wonder how practical this is for a startup; we need concrete ways to implement checks for that input contract drift before code even hits the repository.

Lalam: If we can operationalize those behavioral signals, it could fundamentally improve our entire development culture by making maintainability a measurable quality, not just an afterthought.

Tom: It sounds like the ultimate implication is that evaluation needs to move beyond just task completion and look toward long-term code stability and contract integrity for this kind of work.

Jane: That’s the big picture; we need tools that check for consistency across successive edits, which is a much more mature way to assess agent output than just checking a test pass rate.

Lu: I think the paper's suggestion to use LLM-as-a-judge to track that PR1 behavioral drift is where the creative potential lies for building these better diagnostic tools.

Meng: If we can get those specific drift labels in place, it gives us a concrete way to flag risky code before it gets integrated into larger systems, which is something I can actually build on.

Lalam: That kind of proactive monitoring would be a huge cultural shift, encouraging developers to think about the long-term health of the codebase rather than just getting the immediate fix done.

Tom: Fantastic stuff; so "Is Agent Code Less Maintainable Than Human Code?" isn't just a theoretical exercise; it’s giving us a roadmap for building smarter, more resilient agent code in the long run.

Jane: It really shows that as AI coding tools get more powerful, our quality control and evaluation methods have to evolve right along with them.

Lu: We're still seeing so much potential here for how AI can become part of a self-correcting system rather than just a code generator.

Meng: I’m looking forward to seeing how the practical implementation of these behavioral drift labels looks in real-world scenarios soon.

Shaswat Patel, Betty Li Hou, Arun Purohit, Kai Xu, Jane Pan, He He, Valerie Chen

New York University · Carnegie Mellon University

cs.SE, cs.AI, cs.CL

Submitted: 2026-06-19

Updated: 2026-09-29

Code: https://github.com/shaswatpatel123/CodeThread

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 89/100

The gist: Building on agent code more often lowers downstream resolve rates than building on human code, with effects varying by model and task type.

Key concepts

CodeThread Framework
This framework structures software engineering problems into two sequential steps: an initial implementation task (PR0 to PR1) followed by a dependent follow-on issue resolution (PR1 to PR2). It allows researchers to isolate how the authorship of the first code submission affects the performance of subsequent agent work.
AA vs. HA Scenarios
This compares two specific coding workflows: Human PR1 followed by Agent PR2 (HA) versus Agent PR1 followed by Agent PR2 (AA). Comparing these isolates whether a human-written initial implementation code negatively impacts the performance of subsequent agent work.
Input/Error Contract (IEC)
This concept measures changes in how an agent handles input validation or error messages in its first code submission. When an agent alters this contract, the odds that a human-written initial implementation (HA) will outperform an agent-written one (AA) increase significantly.
PR1 Behavioral Drift Label
This is a label generated by an LLM-as-a-judge to detect subtle changes in input validation or error handling within the agent's first patch. Tracking this drift helps determine if these initial behavioral changes persist into the final code (PR2) and cause test failures.

Terminology

Summary

Building on agent code more often lowers downstream resolve rates than building on human code, with effects varying by model and task type.

CodeThread Framework

CodeThread is introduced as a framework to conduct controlled experiments comparing how well agents can build on and maintain agent versus human code. It transforms standard single-issue software engineering benchmark instances into two-step tasks that comprise an initial implementation task followed by a dependent downstream task. This allows for the isolation of how authorship of the initial task affects downstream agent performance, holding the downstream task author and test-based evaluation fixed. The framework starts from existing issue-resolution coding benchmarks and involves three code states: PR0 (skeleton), PR1 (implementation), and PR2 (follow-on issue resolution).

Experimental Setup

The study applies CodeThread to four frontier models—Claude 4.5 Sonnet, GPT-5, GLM 4.7, and MiniMax M2.5—and on four software engineering benchmarks spanning bug fixing (BF), feature implementation (FI), and refactoring (RF). The authorship scenarios tested are Human PR1 followed by Human PR2 (HH), Human PR1 followed by Agent PR2 (HA), and Agent PR1 followed by Agent PR2 (AA). The comparison between AA and HA isolates the downstream effect of agent-written code on subsequent agent work.

Key Findings on Maintainability

The core finding is that agents more often perform worse when building on agent code than on human code: AA generally underperforms HA, with drops in downstream resolve rate of up to 13.1%. Regression analysis reveals that traditional static maintainability metrics do not explain this difference. Instead, the clearest signals are subtler behavioral differences in agent code, such as changes to input validation and error handling, along with differences in downstream code size and task difficulty.

Sources of Performance Differences

The study identifies three statistically significant predictors of the discordance direction between HA and AA:

  1. Input/Error Contract (IEC): When agent PR1 alters input-validation or error-handling behavior, the odds of HA winning are 1.83× that of AA.

  2. Downstream Code Size Change: A feature set measuring PR2 patch localization shows that AA over-editing at PR2 relative to HA, HA is 1.88× more likely to win over AA.

  3. Instance Difficulty: The instance difficulty score signifies that easier instances each raise the odds that only HA resolves.

Conclusion and Implications

The findings suggest that coding agents introduce maintainability costs that compound as their code is reused or extended in ways current benchmarks and metrics fail to capture. The paper concludes that evaluation must shift from isolated task completion toward long-term maintainability, highlighting the need for safeguards like better review practices to support informed decisions about when and how to rely on agent code. This work points to potential sources of downstream errors introduced by agent code.

The gist: building on agent code more often lowers downstream resolve rates than building on human code, with effects varying by model and task type.

How it works

  1. CodeThread transforms single-issue tasks into two-step chains: an Implementation Task (PR0 → PR1) followed by a Follow-On Issue (PR1 → PR2).

  2. Authorship conditions are set to compare Human PR1 followed by Agent PR2 (HA) against Agent PR1 followed by Agent PR2 (AA).

  3. The final code output, PR2, is evaluated based on whether it successfully resolves the original benchmark issue using the official test suite.

Key Metrics Analyzed

The analysis uses four static proxies for maintenance cost: Cyclomatic Complexity (CC), Cognitive Complexity (CogC), Halstead Volume (HV), and Logical Lines of Code (LLOC). These are computed across supported languages using various open-source libraries and custom tools. Additionally, the study measures patch localization features, such as Jaccard overlaps on edited files and functions between HA and AA PR2 patches.

Behavioral Drift Identification

The paper introduces a PR1 behavioral drift label captured by an LLM-as-a-judge to detect changes in input contract or error handling behavior in the agent's first patch. This divergence is then tracked to determine if it survives into PR2’s final state and whether it directly causes the failing test.

Modeling the Drivers

A logistic regression model is fitted on discordant instances using twelve features, including static maintainability metrics, patch localization features, PR1 behavioral drift labels (IEC), and an instance difficulty score. The model identifies IEC drift in agent PR1, AA over-editing at PR2 relative to HA, and easier instances as the three significant predictors tilting the outcome toward HA resolution.

Improvements for AI systems

Based on the findings of the paper Is Agent Code Less Maintainable Than Human Code?, here are specific, actionable improvements for AI systems and what those improved systems can achieve:


Improvement 1: Implement a Dynamic Maintainability Gate (DMG) in Agent Workflows.

The core finding is that agent code introduces subtle behavioral drift (input validation/error handling) that compounds when built upon by other agents, leading to a significant drop in downstream task resolve rates (up to 13.1%).

The improved system should incorporate a pre-deployment maintainability check during the agent's own development phase.

The improved AI system can:

  1. Monitor its generated patch (PR1) for high-risk behavioral changes identified by the Input/Error Contract (IEC) indicator.

  2. Automatically trigger a secondary, lightweight validation pass specifically targeting input validation and error handling logic, rather than just functional test passing.

  3. If significant deviation from established patterns (as captured by the PR1 Behavioral Drift taxonomy) is detected, the system should halt or flag the patch for mandatory human review before it is committed to a shared base codebase.

Improvement 2: Integrate Context-Aware Code Reuse Scoring for Downstream Tasks.

The study shows that agent code often performs worse than human code when used as a base for subsequent tasks, particularly when those tasks involve larger edits or higher task difficulty (e.g., Refactoring or FI+RF).

The improved AI system should maintain a Code Maintainability Score (CMS) for every component it generates.

The improved AI system can:

  1. When generating new code, it must analyze the existing codebase context to estimate the potential downstream impact of its own output on future modifications.

  2. It should prioritize generating code that exhibits lower PR1 Behavioral Drift and smaller changes in structural complexity (CC, CogC) and verbosity (HV/LLOC) compared to established human code patterns for that specific module.

  3. When tasked with a new dependent task, the system should explicitly query the CMS of its inherited code base and adjust its internal planning to favor modifications that minimize the risk of triggering the observed AA underperforms HA scenario.

Improvement 3: Shift Evaluation Metrics from Static Proxies to Behavioral Contract Verification.

The paper proves that traditional static metrics (CC, Halstead Volume) are insufficient because they fail to capture subtle behavioral differences like input gating or exception handling. The most significant predictor of failure is the Input/Error Contract (IEC) indicator and changes in PR1 behavior.

The improved AI system should prioritize these behavioral signals over purely structural ones during self-evaluation.

The improved AI system can:

  1. Instead of solely relying on static complexity scores, it must include a mandatory check for Behavioral Divergence categories (e.g., Input Gating, Silent Fallback) in its self-test suite.

  2. It should use an LLM-as-a-judge to explicitly verify the integrity of input/output contracts between its intermediate states (PR1 and PR2), specifically looking for IEC drift.

  3. The system's success metric for a task should be weighted not just on test pass rate, but on the stability of the code contract across successive edits, rewarding solutions that maintain consistency with established error-handling patterns.

Improvement 4: Develop Long-Horizon Trajectory Resilience Training.

The current study is limited to two-step chains (PR1 → PR2). The paper suggests that behavioral drift can persist and compound over long development trajectories.

The improved AI system should be trained or designed to anticipate compounding errors across multiple sequential tasks.

The improved AI system can:

  1. Be trained on synthetic long-horizon benchmarks (like CodeFlowBench) that require chaining many dependent fixes, rather than isolated issues.

  2. Implement a mechanism that tracks the cumulative PR1 Behavioral Drift of its generated code across successive edits within a single project context.

  3. If the cumulative drift exceeds a predefined threshold, the system should proactively suggest refactoring or code consolidation to prevent compounding maintainability costs before it is integrated into larger systems.

Improvement 5: Enhance Human-in-the-Loop Feedback for Maintainability Signals.

Since the paper suggests that subtle signals are often missed by metrics, human feedback is crucial for calibrating the system's understanding of maintainable code in a specific context.

The improved AI system should actively solicit and integrate human feedback specifically on maintainability aspects, not just functional correctness.

The improved AI system can:

  1. When encountering ambiguous or complex code structures, it should proactively generate alternative implementations and ask a human reviewer (or an LLM-as-a-judge trained in maintainability) to score the proposed change based on factors like readability and error handling robustness, rather than just functional correctness.

  2. It should use the results of these maintainability evaluations to fine-tune its internal models, explicitly teaching it which types of PR1 changes correlate with better downstream performance in real-world scenarios.

Abstract

Maintainability is a core dimension of software engineering, shaping how code is written, reviewed, and developed over time. While coding agents have demonstrated strong performance on single-issue tasks, it remains unclear how maintainable their code is when future agents build on top of it, potentially leading to compounding downstream effects. We investigate how agent code compares to human code in these maintenance settings, presenting CodeThread, a framework to construct controlled experiments from repository-level coding benchmarks. Applying CodeThread to four frontier coding agents and four benchmarks, we find that agents are less effective at resolving tasks when building on agent code compared to human code, with task resolve rate drops of up to 13.1%. Regression analysis reveals that many traditional software engineering maintainability metrics do not explain this difference. Instead, the clearest signals are subtler behavioral differences in agent code, such as changes to input validation and error handling, along with differences in downstream code size and task difficulty. These findings highlight the need to evaluate these systems not only by immediate task resolution but also by code maintainability, and point to potential sources of downstream errors introduced by agent code.

Sources

Related papers