Professional Software Developers Don't Vibe, They Control: AI Agent Use for Coding in 2025
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Professional Software Developers Don't Vibe, They Control: AI Agent Use for Coding in 2025".
Jane: The paper was written by Ruanqianqian Huang, Avery Reyna, Sorin Lerner, Haijun Xia and Brian Hempel from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Paper summary — Tom, Jane, Lu, Meng, Lalam discuss the paper 'Professional Software Developers Don't Vibe, They Control: AI Agent Use for Coding in 2025' — the thesis, the key findings and why it matters.: Tom: So we started by summarizing the core argument of "Professional Software Developers Don't Vibe, They Control: AI Agent Use for Coding in two thousand twenty-five" and if I understood it right, the central thesis is that developers are moving into a mode of active governance over AI tools.
Jane: Exactly. The paper argues that the relationship isn't one of passive suggestion; it's about deliberate, controlled orchestration of powerful AI agents to build software systems.
Lu: What really caught my attention was how it differentiates between mere assistance and true control—it's a huge conceptual leap for the industry.
Meng: And that difference has massive implications for testing and validation. If you can't trace the agent's reasoning, you can't trust the output, no matter how clean it looks.
Lalam: The implication for culture is huge; it moves us away from a 'code-first' mindset toward a 'verification-first' mindset, which changes everything about quality assurance.
Tom: It feels like we’re moving past the idea of the AI being the primary coder and back into the developer being the conductor of an incredibly powerful orchestra.
Jane: But that orchestration requires developers to develop entirely new skills, doesn't it? It's not enough just knowing Python or Java anymore; they need agent management skills.
Lu: I think we should really focus on what *kind* of control they need to exert—is it prompt engineering, or something deeper in the execution environment?
Meng: Honestly, I suspect it requires a better abstraction layer that lets the developer define constraints and guardrails for the agents, rather than just writing long prompts.
Lalam: Building on Meng’s point, if we can design systems that *force* transparency in agent decisions, we unlock levels of trust and reliability previously unimaginable in complex software.
Tom: It sounds like the paper is setting out a really high bar for what professional development looks like in two thousand twenty-five and beyond.
Jane: And it emphasizes that this control aspect is what makes the difference between functional code and robust, maintainable, enterprise-grade systems.
Page 1 of the paper — Discuss page 1 of the paper: what is new on this page, explained in simple terms. Do not repeat the thesis already covered; build on it.: Tom: So we’re diving into Page one of "Professional Software Developers Don't Vibe, They Control: AI Agent Use for Coding in two thousand twenty-five" and if we established that developers are taking control, what does this page specifically add to that picture?
Jane: This section seems to be setting the historical context, moving away from early LLM use cases and pinpointing the moment when the agent capability becomes truly structural.
Lu: It really highlights how current AI tools are often just sophisticated wrappers around existing workflows, but they aren't fundamentally changing the *developer role* yet.
Meng: What I found interesting on Page one is the discussion around complexity scaling; as systems get bigger, the manual effort required for oversight grows exponentially, which is where agents supposedly help.
Lalam: The page seems to be pointing out that the sheer volume and interconnectedness of modern software demands a higher level of systemic intelligence than just good coding practices alone can provide.
Tom: So, it's not about making individual functions better; it's about managing the *interactions* between many components that an agent might touch.
Jane: Right. It seems to be arguing that the limitations of current AI are often tied to their inability to maintain long-term, multi-step context across an entire codebase.
Lu: This suggests we need agents with persistent memory and a robust internal model of the entire project structure, not just the file they're currently looking at.
Meng: And from an implementation standpoint, that means the
Page 2 of the paper: Tom: So what these developer quotes really hammer home is that just telling an AI to "make a feature" isn't enough; it’s about giving it a whole development protocol, like an actual software life cycle.
Jane: Exactly, Tom. It sounds like the biggest shift isn't in the intelligence of the models themselves, but in how we have to structure our requests for them—it’s almost like writing a detailed instruction manual *for* the AI to follow while it codes.
Meng: That makes sense from an engineering standpoint; if you don't define where variables live or what files need modifying, the code is just a collection of random chunks that won't compile together. You need file paths and function signatures specified upfront.
Lu: And that specificity goes beyond just structure; it touches on the meta-level of thinking—the AI needs to understand *why* we are building something, not just *how* to write the lines of code based on a single prompt.
Lalam: Thinking about that process gives me pause because it suggests that the value shifts from pure coding ability to superior system design and context management, which is profoundly human.
Tom: Right, Lu's point hits on something huge; it means we have to be expert prompters who are also expert architects. Jane, when you hear them talk about "HITL"—human in the loop—what does that actually mean for a team day-to-day?
Jane: It means the AI can write the draft code super fast, but a human still has to jump in after each major task to check for drift or subtle logic errors that only an experienced eye would catch.
Meng: From my side, it translates into needing better tooling—the AI shouldn't just spit out code; it should submit pull requests that are already scaffolded with tests and clearly labeled with which part of the original PRD they fulfilled.
Lu: If we could build a system that dynamically updated its own understanding of the project context based on every commit and every PR, we might finally move past these manual "human in the loop" checks.
Lalam: The cultural impact there is immense; if agents can handle most of the rote iteration and error-checking, developers gain massive bandwidth to focus exclusively on those genuinely novel, high-level architectural decisions that move humanity forward.
Tom: So it’s less about writing code and more about guiding the *system* that writes the code. It feels like we're entering an era where documentation is just as critical as the source code itself.
Jane: It really changes the role of junior developers, too; they won't be learning syntax so much as they'll be learning how to ask incredibly precise questions.
Meng: And that means our training programs need a massive overhaul to teach people how to think like prompt engineers and system integrators.
Lu: I think we might even see a formal academic field emerging just for "Agent Orchestration Protocols," because the complexity of managing these multi-stage, context-heavy workflows is staggering.
Lalam: Ultimately, this advancement doesn't just speed up coding; it elevates the collective human capacity for complex thought by automating away the mundane friction points of development, allowing us to tackle problems previously deemed too massive. So if we nail down these protocols and treat AI agents like highly capable but context-starved junior developers who need constant direction, where do you think this whole process is going next?
Page 3 of the paper: Tom: So, going over these survey responses really drills home a point; it’s less about having a magic prompt and way more about building an entire scaffolding around the AI agent's work. Jane, what’s the simplest way to explain this shift in developer mindset?
Jane: It sounds like developers aren't just asking for answers anymore; they're acting like project managers who are guiding a very smart, but totally fresh-faced, intern through a complex build. They have to give the context upfront.
Meng: That rings true from an engineering standpoint because if you don’t tell the agent exactly which files it needs to look at—the function signatures, the variable names—it's going to guess, and guessing code breaks everything in a large repository.
Lu: But this idea of "controlling" the agent suggests that we're moving beyond just code generation into true workflow orchestration; imagine if an agent could manage the entire Software Development Life Cycle itself!
Lalam: Exactly, Lu points out that this emphasis on process is huge for culture because it means we’re building tools that require *thought*—they need the user to think about their architecture and planning before they even press 'run'.
Tom: And those developers are talking about providing multiple artifacts—the PRD, the architecture plan, maybe even a todo list—before they ask for the first line of code. It’s like building a blueprint before laying the foundation.
Jane: So, if I understand this right, it means that simply saying "make this feature" isn't enough; you have to hand it a whole package deal with rules and context attached.
Meng: That's critical because referencing specific file paths or even recent commits gives the agent immediate understanding of the *current* state of the project, which saves so much time in debugging drift.
Lu: It makes me think about how we need to build meta-agents that specialize only in context management, something that keeps track of all those interconnected files and architectural decisions over months.
Lalam: And that capability shift impacts how teams collaborate; if the AI can reliably maintain context across a large team's contributions, it drastically reduces communication overhead and boosts collective focus.
Tom: It really suggests that the future success of these agents hinges on our ability to structure our requirements—we have to be extremely precise with our own internal processes.
Jane: So, instead of just giving a task, you’re giving a detailed set of instructions that tells the AI exactly what role it's playing and what guardrails it needs to follow.
Meng: Which means we need to integrate these advanced prompt structuring techniques directly into our IDEs so developers don't have to remember all those complex protocols.
Lu: I wonder if we could develop a standard, universal protocol layer that sits between the developer and the agent, enforcing this structured input regardless of the underlying framework.
Lalam: That universal layer would normalize human intent, translating complex human planning into machine-readable control signals, which is a massive step toward truly autonomous development cycles.
Tom: Wow, so we're talking about turning vague ideas into tightly controlled execution plans—that’s a huge leap for the industry! But what does this mean for the next generation of AI tools?
Page 4 of the paper: Tom: So, looking at this survey data from Page ten really nails down that developers aren't just handing off tasks; they're acting like highly detailed project managers for these AI agents.
Jane: Right? It shows that simply saying, "Build me a login screen," isn't nearly enough to get good code back. The folks answering need to guide the AI through a whole lifecycle.
Meng: That’s what caught my eye—the emphasis on needing an SDLC protocol, or Software Development Life Cycle protocol, before they even start coding. It suggests that prompt engineering is really just advanced process documentation right now.
Lu: Exactly! It implies that current LLMs are brilliant executors but lack intrinsic adherence to formalized engineering discipline unless you force it upon them via the prompt structure itself.
Tom: You mean developers have to act as the architect *and* the QA tester simultaneously just to get a working prototype? That’s a huge cognitive load for them, isn't it?
Jane: I think it is, Tom. It means that treating AI like a smart kid who knows nothing about your company or codebase is actually spot on; you have to write all the rules for it.
Lalam: Considering what Jane said, this really highlights that the most valuable part of development in two thousand twenty-five won't be coding speed, but context management and rule definition.
Meng: From a practical standpoint, how do we automate providing that deep context—the PRDs, the architecture plans—without making the prompt itself hundreds of pages long? That’s where the bottleneck is.
Lu: Maybe the next generation of agents won't just *write* code; they'll be designed to *ingest and maintain* a structured project knowledge graph derived from those plans, making retrieval seamless.
Tom: So we move past describing the context and start feeding it a live, queryable model of the whole system? That’s a massive leap from just pasting files into the prompt box.
Jane: It sounds like the AI needs to understand relationships between components—how one file affects another—rather than just treating them as isolated pieces of text.
Lalam: If we can make that context graph automatic, it changes everything for human creativity; it frees up our minds from manually documenting every single constraint and allows us to focus purely on the revolutionary idea itself.
Tom: Okay, so the future isn't just about better prompts; it’s about building an entirely new scaffolding *around* the AI that keeps track of all those protocols and relationships automatically.
Jane: That makes me wonder what happens when we get to that point—what does a fully context-aware, protocol-following agent actually build next?
Page 5 of the paper: Tom: So, if we look at what Page thirteen is laying out, it really zeroes in on moving past just giving the AI a simple task description and into building actual development protocols around it.
Jane: Exactly, Tom; it shifts the focus from *what* you want coded to *how* you need to guide the AI through thinking about that code step by step.
Meng: That’s what I found most interesting—it treats the agent not as an oracle, but as a very capable but inexperienced junior developer who needs clear guardrails and multiple checkpoints along the way.
Lu: It suggests that the future isn't just about better models; it’s about building better *scaffolding* around those models, making the interaction itself a structured process.
Lalam: And that structure, as Lalam notes, means we're not just automating code generation; we’re automating the *thought cycle* of engineering itself, which is a massive cultural shift.
Tom: Right, Lu touched on the scaffolding part—it's like they’re saying you can't just hand over a massive project and hope for the best; you have to break it down into tiny, verifiable chunks.
Jane: Think of it like writing an essay where you don't write the whole thing at once, but first draft an outline, then write the introduction based on that outline, and only then tackle the body paragraphs one by one.
Meng: Practically speaking, this means we need better tooling that can manage those dependencies between tasks—if Task B fails because Task A didn't properly finalize its API endpoint structure, the system needs to know exactly where to go back and fix it.
Lu: It really brings up the idea of required human oversight at crucial junctures; they aren't suggesting full autonomy yet, but rather a highly optimized "human-in-the-loop" process integrated into every step.
Lalam: That human element is key because it means the AI becomes an incredible co-pilot for critical thinking, improving our collective ability to verify complex logic and maintain high standards of craftsmanship in software.
Tom: So, if I’m following this thread correctly, the biggest change isn't in the model's intelligence itself, but in how much context and process we force *into* the prompt structure.
Jane: Precisely; it’s about providing that architectural blueprint alongside the coding request every single time.
Meng: It forces us to be more meticulous with our own documentation and planning before even talking to the AI, which is actually a good thing for overall development quality.
Lu: It elevates the role of the architect—the person who can define those protocols—to an even higher level of importance than ever before.
Lalam: Knowing this focus on controlled process, I wonder what happens when these structured workflows start applying to non-code domains, like scientific hypothesis generation or complex legal drafting?
Page 6 of the paper: Tom: So, we were talking about how giving detailed prompts is great, but what this new section on page sixteen really hammers home is that just *telling* the AI what to do isn't enough anymore.
Jane: Exactly! It seems like the shift isn't just about better writing prompts; it’s about building whole *systems* around how we use AI for coding.
Lu: I was really struck by the emphasis on formalized protocols, which suggests that treating AI like a black box is totally outdated.
Meng: That makes sense from an engineering standpoint; you can't just hand over a feature request and expect it to run in production without guardrails, right?
Lalam: It points toward a fundamental shift in our creative processes, moving from pure ideation to highly controlled, iterative development cycles supported by AI.
Tom: You hit the nail on the head with guardrails, Meng; it's like they’re saying we need an entire Software Development Life Cycle protocol built into the prompt itself.
Jane: Right? It’s less about writing a single perfect prompt and more about managing a conversation that follows established steps—like having to provide the Product Requirements Document first.
Lu: I wonder if this means future AI agents will actually require their own internal state management, remembering every PRD change and every architectural decision we make along the way.
Meng: Remembering stuff is one thing, Lu, but enforcing *adherence* to those steps—like forcing a human-in-the-loop check after every major component build—that’s the hard engineering part they're focusing on.
Lalam: I think that level of mandated checking actually improves our collective culture; it forces us to slow down and value structured validation over speed alone.
Tom: So, if we boil this down, it’s not just about code completion; it’s about process control—making the AI follow the established rules of how software gets built.
Jane: It sounds like developers are realizing they have to become expert prompt *orchestrators*, guiding the AI through multiple distinct stages rather than just giving one big command.
Lu: And those stages need to be clearly delineated, almost like we’re designing a machine that runs in sequential, verifiable passes.
Meng: If we can automate the context handoff—making sure the AI sees the latest commit *and* the original design spec simultaneously—that’s where massive time savings will come from.
Lalam: That seamless integration of context and process is what allows us to build something that feels genuinely intelligent, something that evolves with human oversight guiding every step.
Tom: Wow, so we're moving from giving directions to building the entire assembly line for development; speaking of assembly lines, next up we’ve got to talk about how these agents interact with existing codebases...
Page 7 of the paper: Tom: So, if we’re taking away one thing from these developer quotes, it's that simply asking AI to "build a feature" isn't enough anymore; they’re talking about building entire development protocols around it.
Jane: Exactly, Tom; it seems like the relationship isn't just user and tool anymore—it’s really a highly structured collaboration where the developer has to become an architectural director for the AI agent itself.
Meng: From an engineering standpoint, that "protocol" part is huge; it implies we can’t just throw a prompt at something and hope for the best, we need defined inputs, like PRDs or architecture diagrams attached right there.
Lu: And if they're defining those protocols so rigorously—thinking about SDLC cycles and HITL checkpoints—it suggests that the next frontier isn't just code generation, but autonomous project *management* guided by AI agents.
Lalam: When you think about it culturally, this means that the skill set valued in software development is shifting away from rote coding knowledge toward expert systems thinking and unparalleled clarity in documentation.
Tom: Speaking of clarity, Jane, did those quotes mention anything specific about iteration? It sounded like developers weren't getting it right on the first try, even with great prompts.
Jane: Right, I remember reading that one quote about needing "more guidelines," which tells me that even when the initial context is solid, the AI still requires course correction and refinement to nail complex logic.
Meng: That makes sense; you can give it a file structure, but if the underlying business logic shifts slightly during testing, the agent needs to understand *why* that change is necessary, not just *that* it's necessary.
Lu: Precisely! It’s about feeding it the failure state—the error message and the context around that error—and having it self-correct through a loop of diagnosis and proposal.
Lalam: I think that continuous feedback loop is where the culture changes; instead of seeing bugs as personal failures, we start viewing them as necessary training data points for our advanced AI collaborators.
Tom: So, it's not just about writing good prompts once; it's about maintaining a living document of how the AI should operate on a project over weeks or months?
Jane: Yeah, like having a central "how-to" guide that every developer—and every agent—has to reference before touching a line of code.
Meng: I wonder if that centralization means we’re going to see specialized, highly governed agent frameworks becoming standard tooling rather than just novel research demos.
Lu: Absolutely; think about how we'll need specialized orchestration layers just to manage the context switching between those different development phases they mentioned.
Lalam: Honestly, that focus on structure is fantastic for society because it means that complex, multi-stage projects become less dependent on any single genius human mind and more reliant on defined, auditable processes.
Tom: It really boils down to turning creative guesswork into engineered repeatability, doesn't it? Speaking of repeatability, I wonder what happens when we move from guiding agents through coding protocols to guiding them through entire product launches...
Page 8 of the paper: Tom: So, if I'm wrapping up what we covered earlier, it seems that successful interaction with AI coding agents isn't really about writing one perfect prompt; it’s more about establishing a whole development *protocol* for them to follow.
Jane: Exactly, Tom. It feels like the barrier is shifting from "Can the AI code this?" to "How do we best structure the process so the AI adheres to our standards while coding it?"
Lu: I find that really fascinating because it implies a massive leap in trust—we’re not just asking for an answer; we’re outsourcing adherence to a multi-stage plan, which is way beyond simple instruction following.
Meng: Adherence is the word right there, Lu. From an engineering standpoint, this suggests that simply having a prompt isn't enough; you need tooling that can enforce context boundaries and track dependencies across the whole project structure automatically.
Jane: You know, when people hear "SDLC protocol," they might tune out, but what I hear is this: we're teaching the AI how *we* build things in real company environments, step by step.
Tom: It’s like going from telling an intern to "make me a website" to giving them a detailed project plan that includes wireframes, backend APIs, and QA testing phases.
Lalam: And what's beautiful about that level of structure is how it forces transparency; the agent has to show its work at every single stage, which builds incredible reliability into the output.
Meng: Reliability is key; if we can mandate human-in-the-loop checks after major components, that drastically reduces the risk associated with letting AI build massive, interconnected systems unsupervised.
Lu: But I wonder if this focus on strict protocols might accidentally stifle some of the truly novel, boundary-pushing ideas that come from less structured brainstorming sessions?
Jane: Maybe not so much; perhaps the structure itself becomes a tool for discovery, guiding us through the known possibilities until we hit something unexpected.
Tom: So it's becoming this highly contextual, almost collaborative environment where we're basically building a digital co-pilot that understands our entire corporate memory.
Lalam: That understanding of history and context is what truly elevates the experience; it means the AI isn't just writing code for today’s problem but is contributing to the long-term, evolving culture of our software engineering practice.
Meng: Right, so we need to start thinking about how these protocols integrate into version control systems, making the context part of the commit history itself.
Tom: That makes sense—it’s about keeping the *why* alongside the *what* when we save a change. This focus on rigorous process definitely changes how we evaluate AI assistance moving forward, doesn't it? Speaking of structure, I bet that leads us perfectly into how these agents handle complex data interactions...
Page 9 of the paper: Tom: Page twenty-five really summarizes some of the major themes we've been discussing, particularly in how developers view their relationship with AI agents—it’s definitely not a simple "vibe" situation.
Jane: What I found so compelling here is that the results show a clear division in agent suitability: they are great for small, repetitive tasks but struggle immensely when complexity or domain knowledge increases.
Meng: That difficulty with complexity is the core practical challenge; if an agent can't handle general debugging or large-scale refactoring reliably, it doesn't save us time—it just moves the problem to a different stage.
Lu: But I’m interested in the "controversial" tasks, like high-level planning and brainstorming; those are the areas where human intuition and AI assistance seem to be in a constant, productive tension.
Lalam: From a cultural standpoint, that tension is beautiful because it means the AI isn't trying to replace us; it's acting as an incredibly powerful sparring partner, challenging our blind spots and forcing deeper thought.
Tom: So even when they use agents for planning, they are still heavily involved in the back-and-forth, not just delegating a massive task and hoping for the best.
Jane: Yes, the paper makes it clear that no one in their sample felt comfortable letting an AI handle something critical without a human review or override.
Meng: That ties directly into the issue of business logic; if an agent can't handle domain-specific rules—the things unique to our company—it’s just generating technically correct but functionally useless code.
Lu: The inability to master complex, proprietary knowledge is a huge constraint on current LLMs, and that's where the theoretical research needs to go next: how do we teach the AI that private context?
Lalam: If we solve the context problem, then the cultural shift becomes monumental; we move from just optimizing existing workflows to enabling entirely new classes of creative problem-solving.
Tom: It sounds like there's a huge consensus that AI is not yet a replacement for human expertise, but rather a powerful accelerator that requires expert supervision.
Jane: And this leads us to the next big question: if they are so good at scaffolding and simple fixes, what about the big-picture strategic tasks?
Meng: The data suggests agents are really bad at high-stakes scenarios—security or deployment—which is a massive red flag for enterprise adoption.
Lu: I think that highlights a crucial difference between academic LLM benchmarks and real-world, messy corporate software where things are never perfect.
Lalam: The implication of this constraint is that we must redefine "high performance" to include human oversight, not just code speed. So, how do these findings influence the future development of AI tools?
Conclusion: Tom: So, if I’m summing up what this paper really hammered home for us today, it seems that just having AI write code isn't enough; the real superpower is how developers are structuring their *control* over those agents.
Jane: Exactly, Tom. It shifts the focus from "Can AI write it?" to "How well can I guide the AI through a complex process?" It’s less about magic and more about methodology.
Meng: From an engineering standpoint, that rings true; we always knew that detailed planning—like having a full PRD before touching the first line—is non-negotiable, and this just validates that for agent workflows.
Lu: But think about it, Meng; if we nail the context setting as much as the paper suggests, those planned steps could unlock entirely new kinds of rapid prototyping we haven't even dreamed of yet.
Lalam: It suggests a fundamental change in how intellectual work is measured—it’s moving away from pure output volume and toward the quality of structured oversight, which is fascinating for culture.
Tom: That idea about structured oversight really hits home when you look at those detailed prompt techniques the respondents were using; they weren't just asking for a feature, they were providing an entire development lifecycle within the prompt itself!
Jane: It makes it sound like good old-fashioned software development life cycle discipline, but now you’re doing it conversationally with an AI. You have to be meticulous about defining the boundaries.
Meng: And that meticulousness isn't just writing better prompts; it's about creating reusable scaffolding—the architectural context files, the style guides—that the agent can reference repeatedly without getting confused.
Lu: Precisely! Imagine feeding an agent not just code, but an entire repository history and commit message style guide simultaneously; that level of contextual richness is where the real creativity explodes.
Lalam: Because when you provide that rich, controlled context, the AI stops acting like a general knowledge bot and starts acting like a highly specialized, tireless junior team member who just needs clear direction on company culture and coding standards.
Tom: So, while we're definitely thrilled about how agents can become these hyper-efficient pair programmers by two thousand twenty-five it boils down to us still being the conductors of the orchestra.
Jane: We’ve seen that control is everything—the plan, the structure, the iterative feedback loop—that’s what makes all the difference between a cool experiment and actual professional software.
Meng: I guess for us right now, it means we need to get much better at prompt engineering as a core deliverable skill on our resumes.
Lu: And I can't wait to see how this principle of "control" applies when we move into multimodal agents next year; the planning scaffolds will have to become even more complex!
Lalam: It shifts the focus of human genius from raw execution power to superior systemic design, which really elevates the value of human thought in a cultural sense.
Tom: Well, this has been an incredible deep dive into professional agent use for coding; it certainly gives us a lot to chew on for next time.
Jane: We really appreciate you walking us through how much structure goes into making these advanced tools actually work day-to-day.
Meng: Thanks for the breakdown; I feel like I have a clearer roadmap for what practical implementation needs to look like now.
Lu: You've given me some fantastic ideas about extending this control framework to other complex engineering domains, too.
Lalam: We’re genuinely excited about how these controls will reshape human interaction with technology in the coming years.
Tom: And that leaves us with a whole new set of theories waiting for us, because next up, we're shifting gears completely and talking about generative models for biology...
Ruanqianqian Huang, Avery Reyna, Sorin Lerner, Haijun Xia, Brian Hempel
cs.SE, cs.AI, cs.HC
Submitted: 2026-08-18
Updated: 2026-08-20
Code: https://github.com/Significant-Gravitas/AutoGPT
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 89/100
The gist: * This paper investigates "how experienced developers use agents in building software, including their motivations, strategies, task suitability, and sentiments." The study employed a two-part
Key concepts
- Active Governance
- Developers are shifting from passive suggestion to actively governing powerful AI agents. This means deliberately orchestrating agents to build software systems, requiring new skills beyond basic coding knowledge.
- Verification-First Mindset
- The paper suggests a shift away from a 'code-first' approach toward a 'verification-first' mindset. This emphasizes that tracing an agent's reasoning is crucial for trust and quality assurance in software development.
- Agent Orchestration Protocols
- This refers to the need for structured input, like detailed instructions or SDLC protocols, to guide AI agents. This moves beyond simple prompting into defining precise rules and guardrails for how the agent should execute tasks.
- Context Management
- The future involves moving beyond simply pasting files into prompts. The goal is to build systems where agents can ingest and maintain a live, queryable knowledge graph of the entire project, understanding relationships between components automatically.
Terminology
Summary
This paper investigates how experienced developers use agents in building software, including their motivations, strategies, task suitability, and sentiments.
The study employed a two-part methodology consisting of 13 field observations (P-participants) and a qualitative survey of 99 experienced developers (S-participants). The study sought to answer four research questions (RQs): Motivations (RQ1), Strategies for use (RQ2), Task Suitability (RQ3), and Sentiments (RQ4).
Motivations for Using Agents in Software Development (RQ1)
Experienced developers utilize agentic tools primarily to achieve two core goals: personal productivity and the maintenance of software quality.
-
Productivity: Developers valued agentic tools for
improving the speed of software development.
Many reported a significant boost, with one participant stating,I think my productivity has increased ten-fold.
Factors influencing tool selection included ease of discovery and low barrier to entry, integration into existing tools, and the availability at low or no cost. -
Software Quality: Developers also valued
many software quality attributes,
such as correctness and readability.
Strategies for Using Agents in Software Development (RQ2)
The study found that experienced developers do not employ a vibe coding
approach; instead, they maintain strict control over the agent’s output through planning, supervision, and specific prompting techniques.
-
Controlling Design and Implementation: In all 13 observations, participants
controlled the design of new features to be implemented,
regardless of their familiarity with the task. They ensured that agents work on only a few tasks at a time (averaging 2.1 steps per prompt). -
Prompting Strategies: Developers employed sophisticated prompting techniques, including using
clear context and explicit instructions
and incorporating various types of context such asUI or Design Term,
Reference to Input File,
and Specific Library or API. -
Verification Tactics: To ensure quality, participants used verification strategies like
manual testing, iterating with agent
(P3, P7) orreview[ing] code line-by-line.
Agentic Task Suitability (RQ3)
The suitability of agents varies significantly depending on the complexity and nature of the tasks.
-
Suitable Tasks: Agents are highly effective at accelerating
straightforward, repetitive, and scaffolding tasks,
including: -
Generating new code or boilerplate.
-
Writing tests (19:2).
-
Refactoring (general) or code improvement.
-
Writing/updating documentation (20:0).
-
Simple debugging or fixing simple fixes (12:3).
-
Prototyping and exploring alternatives.
-
Unsuitable Tasks: As task complexity increases, agent suitability decreases. Developers found agents were
unsuitable for business logic or tasks requiring domain knowledge.
They also avoided agents for high-stakes or privacy-sensitive tasks, complex logic, and core architectural changes.
Sentiments Towards Agents (RQ4)
Overall, the sentiments are overwhelmingly positive, provided the developer remains in control.
-
Enjoyment: Participants reported a high level of enjoyment (5.11/6 average rating) compared to working without agents. Some found coding
fun again,
stating that AI has made themless stressed.
-
Control vs. Delegation: While many use agents for collaboration—viewing them as a means of
elucidating thought processes
or acting as arubber ducky
—they never allow the agent to be completely autonomous. The participants are activelyreading the output and steering,
maintaining human oversight even when using tools to generate code.
Conclusion
The research concludes that experienced developers do not currently engage in vibe coding.
Instead, they utilize AI agents as a powerful source of collaboration while strictly adhering to software engineering principles and active supervision. The findings suggest that while agents are excellent for accelerating straightforward tasks, their utility diminishes with complexity or when replacing human expertise.
Improvements for AI systems
As a fastidious AI researcher, I have analyzed the findings of this paper to identify critical gaps in current agentic systems. The core failure mode is a lack of human-centric control and adherence to software engineering rigor, leading to vibes
rather than verifiable results.
The following architectural and behavioral improvements are necessary for the next generation of agentic coding systems:
1. Mandatory Multi-Stage Planning Module (Addressing RQ2 & Table 2):
-
Improvement: Implement a mandatory, explicit planning phase before any code generation is executed. The system must generate a structured implementation plan (e.g., in Markdown or JSON format) that includes the necessary steps and corresponding file changes.
-
What it does: This forces the agent to
think aloud
and verify logic against requirements, preventing premature execution and aligning with the expert strategy of creating a draft plan before human revision.
2. Contextual Knowledge Base Integration (Addressing Table 3 & S85):
-
Improvement: Integrate a dynamic, searchable knowledge base (MCP/RAG) that allows the agents to fetch documentation, API specifications, and codebase context before generating code. This must include reference to specific libraries and architectural decisions.
-
What it does: It eliminates reliance on the LLM's internal weights for up-to-date information, ensuring the agent can handle complex tasks involving new dependencies or prevents errors when encountering
uncommon or private libraries.
3. Iterative Chunking Mechanism (Addressing Table 2 & S59):
-
Improvement: The system must be engineered to execute tasks in small, manageable chunks rather than attempting a monolithic solution. The agent should only proceed to the next step upon explicit human confirmation of successful completion of the current step.
-
What it does: This prevents
running away
from instructions and ensures the developer retains control, addressing the finding that successful agents operate on subsets of tasks at a time.
4. Strict Prompt Engineering Protocol (Addressing Table 3 & S75):
-
Improvement: Enforce a required input structure for all prompts, including mandatory fields for:
-
Task Goal: The high-level objective. -
Context Scope: Reference to specific files and architectural context. -
Constraints: Explicit rules (e.g.,Must be Python 3.10,
orDo not modify the existing database schema
). -
Test Cases: Required unit tests that must accompany the code generation. -
What it does: It forces clarity and specificity, transforming vague requests into actionable engineering tasks, maximizing success rates for straightforward tasks (e.g., 80%+ success when provided with clear context).
5. Automated Verification and Regression Testing (Addressing RQ1 & Table 2):
-
Improvement: Integrate automated verification layers into the agent's workflow. After code generation, the agent must automatically run a suite of unit tests, lint checks, and local functional tests before presenting the output to a human reviewer.
-
What it does: This ensures that
code quality attributes
are maintained, preventing hallucinations and catching errors immediately, allowing developers to trust the output after verification.
6. Human-In-The-Loop (HITL) Gatekeeping (Addressing RQ4 & S13):
-
Improvement: A mandatory
Review Gate
must be placed between every stage of code generation and implementation. The agent cannot commit or execute changes without a required human sign-off, even if the initial results appear successful. -
What it does: This prevents the system from becoming fully autonomous, ensuring that
experienced developers are still in control,
which is crucial for maintaining accountability and preventingvibes
from turning into production failures.
7. Refactoring and Debugging Specialization:
-
Improvement: Programmatically categorize tasks as either
Straightforward/Repetitive
(e.g., generating boilerplate, writing tests, documentation) orComplex/Domain-Specific
(e.g., business logic, architecture). The system should automatically route complex tasks to a specialized, highly constrained mode that requires maximum human oversight and iterative refinement. -
What it does: It ensures the agent is utilized for its strengths—accelerating mundane tasks—while avoiding the pitfalls of high-stakes errors on complex logic, maximizing success rates where agents are most suitable.
Abstract
The rise of AI agents is transforming how software can be built. The promise of agents is that developers might write code quicker, delegate multiple tasks to different agents, and even write a full piece of software purely out of natural language. In reality, what roles agents play in professional software development remains in question. This paper investigates how experienced developers use agents in building software, including their motivations, strategies, task suitability, and sentiments. Through field observations (N=13) and qualitative surveys (N=99), we find that while experienced developers value agents as a productivity boost, they retain their agency in software design and implementation out of insistence on fundamental software quality attributes, employing strategies for controlling agent behavior leveraging their expertise. In addition, experienced developers enjoy working with agents as source of collaboration rather than complete delegation given their judgment for task suitability. Our results shed light on the value of software development best practices in effective use of agents, suggest the kinds of tasks for which agents may be suitable, and point towards future opportunities for better agentic interfaces and agentic use guidelines.
Sources
- Conversational AI as a Coding Assistant: Understanding Programmers' Interactions with and Expectations from Large Language Models for Coding
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
- Evaluating Large Language Models Trained on Code
- Exploring the Challenges and Opportunities of AI-assisted Codebase Generation
- Vibe Coding in Practice: Motivations, Challenges, and a Future Outlook -- a Grey Literature Review
- Exploring Student-AI Interactions in Vibe Coding
- From LLMs to LLM-based Agents for Software Engineering: A Survey of Current, Challenges and Future
- Large Language Model-Based Agents for Software Engineering: A Survey
- The Impact of AI on Developer Productivity: Evidence from GitHub Copilot
- Good Vibrations? A Qualitative Study of Co-Creation, Communication, Flow, and Trust in Vibe Coding
- Vibe coding: programming through conversation with artificial intelligence
- Human-In-the-Loop Software Development Agents
- Agentless: Demystifying LLM-based Software Engineering Agents
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties