Humans are Missing from AI Coding Agent Research
summary
In short
The episode discusses a paper titled "Humans are Missing from AI Coding Agent Research." The hosts explain that while research has focused on autonomous coding agents, the real bottleneck is human-agent collaboration. The paper proposes four pillars—task alignment, steerability, verification, and adaptability—and suggests research directions like building user simulators and defining interaction quality metrics.
Key concepts
- Task Alignment
- This pillar focuses on ensuring the agent understands the user's true intent. It involves comparing the user's intent (z-sub-H) with the agent's internal task specification (z-sub-C). A key issue is the 'clarification spiral,' where agents make assumptions that users reject, leading to endless back-and-forth instructions.
- Steerability
- Steerability addresses whether a user can control the agent during a task, not just at the beginning or end. The paper suggests agents should expose 'control points' at decision moments so humans can intervene and guide the agent instead of waiting for a final output.
- Verification
- This pillar concerns the user's ability to check if the agent's output is correct. The hosts note that standard unit tests are often insufficient because when an agent writes both code and tests, checking the tests becomes a lateral shift in cognitive load rather than a reduction.
- Adaptability
- Adaptability refers to the agent's ability to learn over time and remember user preferences across different tasks. The paper highlights this as a gap because most current agents are stateless, treating every task as if it is the first interaction.
Terminology used across episodes
This episode discusses
- Humans are Missing from AI Coding Agent Research · Paper Radio
- Understanding Software Engineering Agents: A Study of Thought-Action-Result Trajectories
- Evaluation of Text Generation: A Survey
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
- CodeT: Code Generation with Generated Tests
- How can we assess human-agent interactions? Case studies in software agent design
- Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
- Teaching Large Language Models to Self-Debug
- Copilot Arena: A Platform for Code LLM Evaluation in the Wild
- EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
- Agentless: Demystifying LLM-based Software Engineering Agents
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- LLM-Based Multi-Agent Systems for Software Engineering: Literature Review, Vision and the Road Ahead
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
- TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Why AI Agents Still Need You: Findings from Developer-Agent Collaborations in the Wild
- ChatDBG: Augmenting Debugging with Large Language Models
- Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study
The paper
Humans are Missing from AI Coding Agent Research · Read on arXiv
Zora Zhiruo Wang, John Yang, Kilian Lieret, Alexa Tartaglini, Valerie Chen, Yuxiang Wei, Zijian Wang, Lingming Zhang, Karthik Narasimhan, Ludwig Schmidt, Graham Neubig, Daniel Fried, Diyi Yang
Carnegie Mellon University · Stanford University · Princeton University · University of Illinois Urbana-Champaign
Recent progress in AI coding agent research has led to rapid improvements in agents' ability to autonomously perform complex software engineering tasks, from editing large codebases to executing long-horizon development workflows. As these systems make strides, however, the primary bottleneck to practical usefulness increasingly shifts away from pure task-solving capability, and toward challenges in how users communicate with, supervise, and trust agents. In this position paper, we argue for a reorientation from autonomous to human-centered coding agents: systems designed not only to complete tasks, but to collaborate effectively with people. We identify four core interaction-level dimensions that characterize the human-agent task-solving loop: task alignment, verifiability, steerability, and adaptability. Finally, we outline concrete research directions to advance these dimensions, including user-involved coding environments, comprehensive verification mechanisms, and principled measures of human-agent interaction quality.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Humans are Missing from AI Coding Agent Research".
Jane: The paper was written by Zora Zhiruo Wang, John Yang, Kilian Lieret, Alexa Tartaglini, Valerie Chen et al. from Carnegie Mellon University and Stanford University and Princeton University and University of Illinois Urbana-Champaign.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back to the show, everyone. Today we’re digging into a paper that’s been making waves in the AI research community, and it’s called “Position: Humans are Missing from AI Coding Agent Research.” Jane, I gotta say, just the title alone got me fired up.
Jane: Oh, absolutely, Tom. It’s one of those papers that makes you stop and go, “Wait, yeah, why haven’t we been talking about this?” The authors are a big collaboration from Carnegie Mellon, Stanford, Princeton, and UIUC. Zora Zhiruo Wang, John Yang, Kilian Lieret, and a whole bunch of others. These are serious names in the coding-agent space.
Tom: And the core argument is pretty bold. They’re saying that the entire field has been obsessed with one thing: making agents that can solve harder and harder coding tasks all by themselves. You know, benchmarks like SWE-bench where the agent gets an issue and just has to fix it solo.
Jane: Right, and they’re not saying that’s useless. But they’re making the case that the real bottleneck for these tools in the real world isn’t whether the agent can write the code. It’s whether a human can actually work with it. Can you tell it what you want? Can you check what it did? Can you steer it when it goes off track?
Tom: Yeah, they put it really well in the paper. They say the bottleneck shifts from “can it work?” to “can humans understand, trust, and work with it?” And that’s a huge reframing.
Jane: It really is. And they back it up with some interesting data. For example, they looked at agent-generated patches on SWE-bench Verified and found they’re consistently longer than the “gold” human-written patches. Sometimes two or three times longer. That makes them harder for a human to review and verify.
Tom: So the agent might pass the test, but the code is bloated and messy. That’s a real problem for a developer who has to maintain that code later. It’s not just about getting a green checkmark on a benchmark.
Jane: Exactly. And that’s why the authors are calling for a reorientation. They want to move from autonomous agents to human-centered agents. Systems designed to collaborate, not just to complete.
Tom: I love that. It’s such a simple idea, but it changes everything about how we measure progress. Instead of just asking “did the agent finish the task?”, we should be asking “how well did the human and agent work together?”
Jane: And that’s the hook that’s going to lead us into the meat of the paper. Because they don’t just stop at the problem statement. They actually propose four specific dimensions that define this human-agent loop. We’ll get into those next.
The Paper's Core Argument: Tom: So, Jane, we’ve established that “Position: Humans are Missing from AI Coding Agent Research” is making a big fuss about the field’s obsession with autonomy. But what’s their actual proposal? What do they want us to do about it?
Jane: Great question. They break it down into four pillars, and they call them the interaction primitives of the human-agent task-solving loop. The first one is task alignment. That’s about whether the agent actually understands what you want it to do.
Tom: Right, and they formalize this. They talk about the user having a true intent, let’s call it z-sub-H, and the agent forming its own internal task spec, z-sub-C. And task alignment is basically the similarity between those two. If you say “build me a portfolio website,” the agent needs to figure out you probably want it dark mode, minimal, and maybe using a specific framework.
Jane: And the paper points out that this is where a lot of frustration comes from. Users get stuck in what they call a “clarification spiral.” The agent makes assumptions, the user rejects the output, the user gives more instructions, the agent makes new assumptions, and on and on it goes.
Tom: Yeah, that sounds exhausting. And the second pillar is steerability. This is about whether the user can control the agent *during* the task, not just at the start or the end.
Jane: Exactly. They argue that agents should expose “control points” at meaningful decision moments. Like, “Hey, I’m about to pick a CSS framework, do you have a preference?” Instead of just running off and doing the whole thing and then presenting you with a done deal.
Tom: So it’s not all-or-nothing. You don’t have to choose between full autonomy and micromanaging every keystroke. The agent should know when to check in.
Jane: Precisely. And the third pillar is verification. This is a big one. It’s about whether the user can actually check if the agent’s output is correct. And the paper makes a really sharp point here: unit tests are the standard in benchmarks, but they’re not always great in the real world.
Tom: Right, because if the agent writes both the code and the tests, the tests aren’t really independent evidence anymore. The user has to check the tests too. That’s a lateral shift in cognitive load, not a reduction.
Jane: And the fourth pillar is adaptability. This is about the agent learning over time. If you tell it once that you prefer logging over print statements, it should remember that next week. The paper calls this out as a major gap because most current agents are stateless. They treat every task like it’s the first time they’ve ever met you.
Tom: So those are the four pillars: task alignment, steerability, verification, and adaptability. And they all feed into each other. If you can’t align on intent, steering is pointless. If you can’t verify, you can’t trust. If you can’t adapt, you’re stuck repeating yourself forever.
Jane: And that’s what makes this paper so compelling. It’s not just a philosophical rant. They give you a framework for thinking about what “good” looks like. And that sets us up perfectly to talk about the concrete research directions they propose. That’s next.
Proposed Improvements: Tom: Alright, so we’ve got the four pillars. But what does the paper actually suggest we *do* about it? I mean, it’s a position paper, but they’re not just sitting on their hands.
Jane: Not at all. They lay out four high-leverage research directions. And the first one is all about scaling human modeling. The problem is that evaluating human-agent collaboration usually means running expensive human studies, which don’t scale, or just defaulting to autonomous benchmarks, which ignore humans entirely.
Tom: And their solution is to build better user simulators. You know, AI models that act like realistic developers. Not the overly-cooperative, homogeneous simulators we have now, but ones with different expertise levels, different preferences, and realistic failure modes.
Jane: Right, they want to mine GitHub data, pull request histories, even real user-agent interaction data, to train these simulators. The goal is to have a diverse cast of virtual users that you can test your agent against, without needing to hire hundreds of human testers every time.
Tom: That’s clever. And the second direction is about enabling efficient oversight. They want the agent to proactively help you verify its own work. Instead of just dumping a diff on you, it should figure out what kind of verification makes sense for the task.
Jane: They have this great vision of “shapeshifting” verification. If you’re building a data pipeline, maybe it shows you a summary of how it handled missing values. If you’re building a webpage, it shows you a rendered preview. If it’s a CLI tool, it records a terminal session so you can watch it run.
Tom: So the agent adapts its output to make it as easy as possible for the human to judge. That’s a big shift from the current “here’s a patch, good luck” approach.
Jane: And the third direction is about defining better measures for interaction. They point out that we have metrics for task success, but almost nothing for interaction quality. How many turns did it take to align on intent? How many interventions did the user have to make? How much effort did verification take?
Tom: And they’re drawing on decades of HCI research for this. They even mention the CUPS taxonomy, which categorizes programmer-AI interactions into twelve states. And they’re saying we should port those ideas into the ML training loop.
Jane: Finally, the fourth direction is to go beyond software engineering. They make the point that coding agents are becoming general-purpose agents. People are using them to manage smart homes, plan weddings, even monitor greenhouses. And in those domains, the four pillars become even more critical.
Tom: Like, if an agent is controlling your front door, you really want to be able to verify what it’s doing before it acts. And if it’s managing a portfolio, “be more conservative” means something different to a retiree than to a college student.
Jane: Exactly. So the paper is pushing us to think about coding agents not as tools for professional developers, but as a new kind of general-purpose assistant. And that’s a pretty exciting future.
Conclusion: Tom: Well, Jane, we’ve covered a lot of ground on “Position: Humans are Missing from AI Coding Agent Research.” Let’s wrap it up for our listeners.
Jane: Absolutely. The paper’s core message is simple but profound. The field has been optimizing for autonomous task completion, but the real bottleneck for practical usefulness is human-agent interaction. They gave us four pillars to think about: task alignment, steerability, verification, and adaptability.
Tom: And they didn’t just stop at the diagnosis. They proposed concrete research directions, like building realistic user simulators, creating adaptive verification mechanisms, defining interaction-quality metrics, and expanding beyond traditional software engineering.
Jane: The big takeaway for me is that we need to stop treating humans as an afterthought. The paper argues that human involvement isn’t a temporary workaround for immature models. It’s intrinsically necessary for accountability, for judgment in novel situations, and for aligning with intent that only humans can ultimately adjudicate.
Tom: That’s a powerful statement. And it’s a call to action for the research community. Are we going to optimize for leaderboards, or for the people who actually use these systems?
Jane: Well said. It’s a question that’s going to shape the next few years of AI research. And with that, we’re going to say goodbye to this paper and get ready to dive into the next one. Thanks for listening, everyone.
Tom: See you on the next episode.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization