Humans are Missing from AI Coding Agent Research
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Humans are Missing from AI Coding Agent Research".
Jane: The paper was written by Zora Zhiruo Wang, John Yang, Kilian Lieret, Alexa Tartaglini, Valerie Chen et al. from Carnegie Mellon University and Stanford University and Princeton University and University of Illinois Urbana-Champaign.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back to the show, everyone. Today we’re digging into a paper that’s been making waves in the AI research community, and it’s called “Position: Humans are Missing from AI Coding Agent Research.” Jane, I gotta say, just the title alone got me fired up.
Jane: Oh, absolutely, Tom. It’s one of those papers that makes you stop and go, “Wait, yeah, why haven’t we been talking about this?” The authors are a big collaboration from Carnegie Mellon, Stanford, Princeton, and UIUC. Zora Zhiruo Wang, John Yang, Kilian Lieret, and a whole bunch of others. These are serious names in the coding-agent space.
Tom: And the core argument is pretty bold. They’re saying that the entire field has been obsessed with one thing: making agents that can solve harder and harder coding tasks all by themselves. You know, benchmarks like SWE-bench where the agent gets an issue and just has to fix it solo.
Jane: Right, and they’re not saying that’s useless. But they’re making the case that the real bottleneck for these tools in the real world isn’t whether the agent can write the code. It’s whether a human can actually work with it. Can you tell it what you want? Can you check what it did? Can you steer it when it goes off track?
Tom: Yeah, they put it really well in the paper. They say the bottleneck shifts from “can it work?” to “can humans understand, trust, and work with it?” And that’s a huge reframing.
Jane: It really is. And they back it up with some interesting data. For example, they looked at agent-generated patches on SWE-bench Verified and found they’re consistently longer than the “gold” human-written patches. Sometimes two or three times longer. That makes them harder for a human to review and verify.
Tom: So the agent might pass the test, but the code is bloated and messy. That’s a real problem for a developer who has to maintain that code later. It’s not just about getting a green checkmark on a benchmark.
Jane: Exactly. And that’s why the authors are calling for a reorientation. They want to move from autonomous agents to human-centered agents. Systems designed to collaborate, not just to complete.
Tom: I love that. It’s such a simple idea, but it changes everything about how we measure progress. Instead of just asking “did the agent finish the task?”, we should be asking “how well did the human and agent work together?”
Jane: And that’s the hook that’s going to lead us into the meat of the paper. Because they don’t just stop at the problem statement. They actually propose four specific dimensions that define this human-agent loop. We’ll get into those next.
The Paper's Core Argument: Tom: So, Jane, we’ve established that “Position: Humans are Missing from AI Coding Agent Research” is making a big fuss about the field’s obsession with autonomy. But what’s their actual proposal? What do they want us to do about it?
Jane: Great question. They break it down into four pillars, and they call them the interaction primitives of the human-agent task-solving loop. The first one is task alignment. That’s about whether the agent actually understands what you want it to do.
Tom: Right, and they formalize this. They talk about the user having a true intent, let’s call it z-sub-H, and the agent forming its own internal task spec, z-sub-C. And task alignment is basically the similarity between those two. If you say “build me a portfolio website,” the agent needs to figure out you probably want it dark mode, minimal, and maybe using a specific framework.
Jane: And the paper points out that this is where a lot of frustration comes from. Users get stuck in what they call a “clarification spiral.” The agent makes assumptions, the user rejects the output, the user gives more instructions, the agent makes new assumptions, and on and on it goes.
Tom: Yeah, that sounds exhausting. And the second pillar is steerability. This is about whether the user can control the agent *during* the task, not just at the start or the end.
Jane: Exactly. They argue that agents should expose “control points” at meaningful decision moments. Like, “Hey, I’m about to pick a CSS framework, do you have a preference?” Instead of just running off and doing the whole thing and then presenting you with a done deal.
Tom: So it’s not all-or-nothing. You don’t have to choose between full autonomy and micromanaging every keystroke. The agent should know when to check in.
Jane: Precisely. And the third pillar is verification. This is a big one. It’s about whether the user can actually check if the agent’s output is correct. And the paper makes a really sharp point here: unit tests are the standard in benchmarks, but they’re not always great in the real world.
Tom: Right, because if the agent writes both the code and the tests, the tests aren’t really independent evidence anymore. The user has to check the tests too. That’s a lateral shift in cognitive load, not a reduction.
Jane: And the fourth pillar is adaptability. This is about the agent learning over time. If you tell it once that you prefer logging over print statements, it should remember that next week. The paper calls this out as a major gap because most current agents are stateless. They treat every task like it’s the first time they’ve ever met you.
Tom: So those are the four pillars: task alignment, steerability, verification, and adaptability. And they all feed into each other. If you can’t align on intent, steering is pointless. If you can’t verify, you can’t trust. If you can’t adapt, you’re stuck repeating yourself forever.
Jane: And that’s what makes this paper so compelling. It’s not just a philosophical rant. They give you a framework for thinking about what “good” looks like. And that sets us up perfectly to talk about the concrete research directions they propose. That’s next.
Proposed Improvements: Tom: Alright, so we’ve got the four pillars. But what does the paper actually suggest we *do* about it? I mean, it’s a position paper, but they’re not just sitting on their hands.
Jane: Not at all. They lay out four high-leverage research directions. And the first one is all about scaling human modeling. The problem is that evaluating human-agent collaboration usually means running expensive human studies, which don’t scale, or just defaulting to autonomous benchmarks, which ignore humans entirely.
Tom: And their solution is to build better user simulators. You know, AI models that act like realistic developers. Not the overly-cooperative, homogeneous simulators we have now, but ones with different expertise levels, different preferences, and realistic failure modes.
Jane: Right, they want to mine GitHub data, pull request histories, even real user-agent interaction data, to train these simulators. The goal is to have a diverse cast of virtual users that you can test your agent against, without needing to hire hundreds of human testers every time.
Tom: That’s clever. And the second direction is about enabling efficient oversight. They want the agent to proactively help you verify its own work. Instead of just dumping a diff on you, it should figure out what kind of verification makes sense for the task.
Jane: They have this great vision of “shapeshifting” verification. If you’re building a data pipeline, maybe it shows you a summary of how it handled missing values. If you’re building a webpage, it shows you a rendered preview. If it’s a CLI tool, it records a terminal session so you can watch it run.
Tom: So the agent adapts its output to make it as easy as possible for the human to judge. That’s a big shift from the current “here’s a patch, good luck” approach.
Jane: And the third direction is about defining better measures for interaction. They point out that we have metrics for task success, but almost nothing for interaction quality. How many turns did it take to align on intent? How many interventions did the user have to make? How much effort did verification take?
Tom: And they’re drawing on decades of HCI research for this. They even mention the CUPS taxonomy, which categorizes programmer-AI interactions into twelve states. And they’re saying we should port those ideas into the ML training loop.
Jane: Finally, the fourth direction is to go beyond software engineering. They make the point that coding agents are becoming general-purpose agents. People are using them to manage smart homes, plan weddings, even monitor greenhouses. And in those domains, the four pillars become even more critical.
Tom: Like, if an agent is controlling your front door, you really want to be able to verify what it’s doing before it acts. And if it’s managing a portfolio, “be more conservative” means something different to a retiree than to a college student.
Jane: Exactly. So the paper is pushing us to think about coding agents not as tools for professional developers, but as a new kind of general-purpose assistant. And that’s a pretty exciting future.
Conclusion: Tom: Well, Jane, we’ve covered a lot of ground on “Position: Humans are Missing from AI Coding Agent Research.” Let’s wrap it up for our listeners.
Jane: Absolutely. The paper’s core message is simple but profound. The field has been optimizing for autonomous task completion, but the real bottleneck for practical usefulness is human-agent interaction. They gave us four pillars to think about: task alignment, steerability, verification, and adaptability.
Tom: And they didn’t just stop at the diagnosis. They proposed concrete research directions, like building realistic user simulators, creating adaptive verification mechanisms, defining interaction-quality metrics, and expanding beyond traditional software engineering.
Jane: The big takeaway for me is that we need to stop treating humans as an afterthought. The paper argues that human involvement isn’t a temporary workaround for immature models. It’s intrinsically necessary for accountability, for judgment in novel situations, and for aligning with intent that only humans can ultimately adjudicate.
Tom: That’s a powerful statement. And it’s a call to action for the research community. Are we going to optimize for leaderboards, or for the people who actually use these systems?
Jane: Well said. It’s a question that’s going to shape the next few years of AI research. And with that, we’re going to say goodbye to this paper and get ready to dive into the next one. Thanks for listening, everyone.
Tom: See you on the next episode.
Zora Zhiruo Wang, John Yang, Kilian Lieret, Alexa Tartaglini, Valerie Chen, Yuxiang Wei, Zijian Wang, Lingming Zhang, Karthik Narasimhan, Ludwig Schmidt, Graham Neubig, Daniel Fried, Diyi Yang
Carnegie Mellon University · Stanford University · Princeton University · University of Illinois Urbana-Champaign
cs.HC, cs.AI, cs.SE
Submitted: 2026-07-04
Code: https://github.com/anthropics/skills
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
Key concepts
- Task Alignment
- This pillar focuses on ensuring the agent understands the user's true intent. It involves comparing the user's intent (z-sub-H) with the agent's internal task specification (z-sub-C). A key issue is the 'clarification spiral,' where agents make assumptions that users reject, leading to endless back-and-forth instructions.
- Steerability
- Steerability addresses whether a user can control the agent during a task, not just at the beginning or end. The paper suggests agents should expose 'control points' at decision moments so humans can intervene and guide the agent instead of waiting for a final output.
- Verification
- This pillar concerns the user's ability to check if the agent's output is correct. The hosts note that standard unit tests are often insufficient because when an agent writes both code and tests, checking the tests becomes a lateral shift in cognitive load rather than a reduction.
- Adaptability
- Adaptability refers to the agent's ability to learn over time and remember user preferences across different tasks. The paper highlights this as a gap because most current agents are stateless, treating every task as if it is the first interaction.
Terminology
Summary
Summary
This position paper argues that AI coding agent research is overly fixated on autonomous task completion, treating benchmark difficulty as a proxy for practical value. The authors hypothesize that the next meaningful breakthroughs lie not in what agents can do solo, but in their interplay with human developers and users.
They contend that as agent capabilities improve, the bottleneck shifts from 'can it work?' to 'can humans understand, trust, and work with it?', making human-centered design essential for translating capability into practical usefulness.
The paper identifies a growing disconnect between agent research and real-world deployment,
noting that much of the field has focused on advancing autonomy, measured by success rate on harder benchmarks,
while real-world programming is rarely a one-shot activity. Instead, it unfolds through iterative interaction, partial delegation, evolving goals, and continuous human oversight.
The authors argue that utility depends not only on whether an agent can complete a task, but on whether humans can effectively work with it to shape outcomes and future work.
The paper's central position is that research should shift from maximizing agent-solo autonomy to maximizing human-centered usefulness.
To make this concrete, the authors identify four key dimensions that characterize human-centered coding agents, which they formalize with measurable definitions:
-
Task Alignment:
the process by which humans and agents establish and maintain a shared task understanding through mutual modeling.
The authors formalize this as the distance between the user-intended specification and the agent's inferred specification. They note thata significant pain point is the 'clarification spiral'
wheremodels launch into implementation without inferring unstated constraints and make incorrect assumptions.
They highlight thatopen conversation data between humans and AI coding systems remains starkly lacking
and thatbenchmarks sidestep the study of such intricacies entirely: pass@k collapses all signal about communicative quality into a single bit.
-
Steerability:
an agent's ability to expose and respond to human control signals throughout task execution.
The authors argue thatpractical steerability requires agents to identify meaningful control points in the task-solving process—moments where alternative execution paths, tradeoffs, or commitments arise—and to expose those choices to the user.
They note thatexisting work is largely descriptive or peripheral
and thatwe lack principled approaches for agents to identify and expose meaningful control points to users, leaving systems trapped between a false dichotomy between full agent autonomy and constant human supervision.
-
Verification:
a user's ability to assess if a coding agent's outputs are correct with respect to task requirements.
The authors note thatusers cannot trust what they cannot verify,
citing thatof the 84% of developers now using AI coding tools, nearly half do not trust outputs while two-thirds report that half-baked solutions lead to heavier debugging burdens.
They argue that "unit testing is the dominant verification paradigm in current benchmarks given their precision and reproducibility. However, in practice, it captures only a narrow notion of functional correctness and often shifts the verification burden onto users.They advocate for verification that
surfaces evidence in human-interpretable ways, such as visual previews and interactive summaries." -
Adaptability:
an AI coding agent's ability to maintain and update itself using accumulated experience, to improve future performance while preserving previously acquired capabilities.
The authors note thata common pain point is the need to re-establish context and re-state execution details every session, often described as 'prompt fatigue'.
They argue thattoday's coding agents are developed and evaluated on isolated, one-off tasks, with no incentive for learning from prior sessions or accumulating user-specific context
and thatwhen adaptability is studied, it optimizes for agents, not users.
The paper outlines four high-leverage research directions to address these gaps:
-
Scale human modeling: Building
executable environments with (simulated) users
anduser simulators that faithfully model humans as active software developers and users.
The authors note thatcurrent simulators prompted to 'act like a human' behave homogeneously and overly cooperative
and thatfaithful simulation requires modeling personal preferences, expertise, and realistic failure modes.
-
Enable efficient oversight: Rather than
placing the full burden on users, coding systems should proactively support oversight.
The authors envisionverification becoming a dynamic, protean procedure
whereconditioned on task and user, agents should reason about notions of quality and surface appropriate artifacts to make validation tractable.
-
Define measures for interaction: The authors note that
the current focus on task resolution leaves interaction quality unmeasured and unoptimized
and thatdecades of HCI and software engineering research offer rich precedent for defining and operationalizing interaction quality from user studies.
-
Go beyond software engineering: The authors argue that
coding agents are general agents
and thatany task expressible programmatically becomes fair game,
citing examples like managing a stock portfolio, controlling a smart home, and tending a greenhouse.
The paper also engages with counter-arguments. Against the claim that AI will take over coding,
the authors argue that whether or not this future materializes, the study of how humans participate and benefit from an AI-enhanced coding loop will remain relevant
and that the locus of effort simply shifts upstream, from implementation to specification and evaluation.
Against the claim that human interaction and evaluation are too costly to scale,
they argue that just as the NLP community developed scalable proxies of human feedback to improve instruction following and alignment,
similar approaches can work for coding agents. Against the capabilities first
argument, they contend that if we continue optimizing solely for autonomous coding agents, we will produce just that. Better collaborators will not emerge for free.
The paper concludes: "We contend that the omission of consideration for humans in AI coding agent research threatens to undermine the very utility these systems seek to provide. Left unchecked, the gap between AI coding agent research and real-world utility may very well widen. The research community must decide whether to optimize for leaderboards or the people who actually use these systems as well. Otherwise, we run the risk of building ever-more-capable coding agents that fewer people know how to wield."
Improvements for AI systems
Based on the paper's core argument, here are specific improvements to AI coding agent systems:
Improvement: Add an explicit intent-inference and clarification mechanism before code generation.
What the improved system can do:
-
Before writing code, the agent generates a structured task specification (z C) and compares it against likely user intent (z H), flagging ambiguities
-
When the instruction is under-specified (e.g.,
make it dark mode
), the agent asks targeted clarifying questions about style, scope, and constraints (e.g.,Dark mode for all pages or just the home page? Should I use CSS variables or inline styles?
) -
The agent surfaces its assumptions explicitly (e.g.,
I'm assuming you want reusable CSS classes rather than Tailwind
) before executing -
It tracks the
clarification spiral
— if the user rejects outputs multiple times, the agent proactively re-asks about intent rather than blindly retrying
Capability Current Systems Improved System
Clarify ambiguous intent before acting Rarely; often guesses Proactively asks targeted questions
Expose decision points mid-task No; runs to completion Surfaces control points with options
Verify in a human-friendly way Unit tests only Task- and user-appropriate artifacts
Remember preferences across sessions No Yes, with conflict detection
Predict user satisfaction in real time No Yes, from interaction traces
Adapt to non-code tasks Limited Full four-pillar support
Abstract
Recent progress in AI coding agent research has led to rapid improvements in agents' ability to autonomously perform complex software engineering tasks, from editing large codebases to executing long-horizon development workflows. As these systems make strides, however, the primary bottleneck to practical usefulness increasingly shifts away from pure task-solving capability, and toward challenges in how users communicate with, supervise, and trust agents. In this position paper, we argue for a reorientation from autonomous to human-centered coding agents: systems designed not only to complete tasks, but to collaborate effectively with people. We identify four core interaction-level dimensions that characterize the human-agent task-solving loop: task alignment, verifiability, steerability, and adaptability. Finally, we outline concrete research directions to advance these dimensions, including user-involved coding environments, comprehensive verification mechanisms, and principled measures of human-agent interaction quality.
Sources
- Understanding Software Engineering Agents: A Study of Thought-Action-Result Trajectories
- Evaluation of Text Generation: A Survey
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
- CodeT: Code Generation with Generated Tests
- How can we assess human-agent interactions? Case studies in software agent design
- Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
- Teaching Large Language Models to Self-Debug
- Copilot Arena: A Platform for Code LLM Evaluation in the Wild
- EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
- Agentless: Demystifying LLM-based Software Engineering Agents
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- LLM-Based Multi-Agent Systems for Software Engineering: Literature Review, Vision and the Road Ahead
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
- TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Why AI Agents Still Need You: Findings from Developer-Agent Collaborations in the Wild
- ChatDBG: Augmenting Debugging with Large Language Models
- Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support