SWE-Together: Evaluating Coding Agents in Interactive User Sessions
summary
The gist
Most coding-agent benchmarks are static, but real coding assistance is interactive, requiring users to clarify goals and correct mistakes over multiple turns.
In short
SWE-Together is a new benchmark that tests coding agents as collaborators by measuring final code correctness and how many times they need corrective feedback during interaction. It reconstructs real user-agent sessions into reproducible tasks, allowing researchers to see how well different models work together in a multi-turn coding environment.
Key concepts
- SWE-Together
- A benchmark built from 11,260 recorded user–agent coding sessions. It tests agents by measuring both the final correctness of their code and the number of corrective feedback turns needed during their interaction.
- Task Construction Pipeline
- A three-step process that transforms raw coding sessions into reproducible tasks. This involves filtering raw data, checking for local reproducibility, and using an LLM judge to confirm if a session can be turned into a verifiable task in an isolated sandbox.
- User Correction
- A diagnostic metric measuring how much the user simulator had to intervene. It tests the hypothesis that stronger models require less steering; lower correction scores suggest better models need less corrective guidance to succeed.
Terminology used across episodes
This episode discusses
- SWE-Together: Evaluating Coding Agents in Interactive User Sessions · Paper Radio
- Program Synthesis with Large Language Models
- SWE-chat: Coding Agent Interactions From Real Users in the Wild
- Evaluating Large Language Models Trained on Code
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants
The paper
SWE-Together: Evaluating Coding Agents in Interactive User Sessions · Read on arXiv
Yifan Wu, Zhuokai Zhao, Songlin Li, Ho Hin Lee, Jiacheng Zhu, Shirley Wu, Tianhe Yu, Serena Li, Lizhu Zhang, Xiangjun Fan
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "SWE-Together: Evaluating Coding Agents in Interactive User Sessions".
Jane: Most coding-agent benchmarks are static, but real coding assistance is interactive, requiring users to clarify goals and correct mistakes over multiple turns.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, let's talk about those authors and what this paper is actually titled. The title of "SWE-Together: Evaluating Coding Agents in Interactive User Sessions" clearly signals the main focus here, which is on the interactive aspect of coding agent performance.
Jane: It’s quite direct, Tom. It tells us immediately that the research isn't about just seeing if an AI can write a piece of code, but how it behaves when things change during the process. That distinction between static and interactive is really important for understanding current capabilities.
Lu: The authors are clearly trying to bridge that gap between what we see in controlled, fixed benchmarks and what happens in real-world software engineering environments where users constantly pivot their needs. It's about capturing that dynamic flow of requirements.
Meng: From an engineering standpoint, the implication is that we need evaluation setups that simulate the actual workflow—where you have to pause and ask questions before moving on. That means testing isn't just running a script; it’s simulating a conversation with a junior developer who keeps changing their mind.
Lalam: I see it as a way to evaluate if an agent is truly becoming a collaborator, not just an autocomplete tool that spits out code based on the first prompt. If they can handle trajectory-conditioned replays, that suggests deeper understanding of the overall engineering goal.
The paper's summary: Tom: So, summarizing what SWE-Together does, it takes those raw, multi-turn coding agent sessions and turns them into reproducible software engineering tasks. They use a three-part pipeline to do this transformation before they even start the actual evaluation process.
Jane: That pipeline is where the heavy lifting happens; first, they filter and normalize those raw sessions to ensure we’re only looking at viable interactions, then they build a sandbox environment with pinned environments and executable checks. It’s a lot of preprocessing work to make sure the tasks are sound.
Lu: The second big piece is this user simulator that replays the original user's intent in a trajectory-conditioned manner, meaning it only steps in when the agent actually needs help based on what happened before, not just on a pre-set schedule. That’s sophisticated simulation design.
Meng: The third part is how they evaluate it: they separate measuring the final code quality from measuring how well the simulator behaves—things like "User Correction" and "Intent Coverage." It shows that correctness isn't the only thing we should care about when we talk about agent performance in real scenarios.
Lalam: That measurement of User Correction is really telling; it suggests that a model doesn't need constant steering to succeed, which points toward a more intuitive understanding of the user’s evolving intent. It’s about achieving alignment naturally over time.
The paper's improvements: Tom: The paper highlights several specific improvements in their methodology, primarily focusing on how they measure agent behavior through diagnostics like Intent Coverage and User Correction. They are looking at how well the simulator matches the original intents and how often users actually need to correct the agent.
Jane: They also introduced a unique way to assess correctness using an agentic rubric judge that gets updated based on running phases, ensuring that even the scoring mechanism is calibrated for fairness across different agents being tested. That makes it more rigorous than just looking at one final output score.
Lu: The paper suggests that we should focus on models that minimize corrective steering because they demonstrate a deeper grasp of the task trajectory. It’s not just about getting the right code; it’s about getting there efficiently with minimal back-and-forth.
Meng: From a practical standpoint, this means we can start optimizing for models that are robust enough to handle ambiguity and provide relevant feedback proactively rather than waiting for a major error. That shifts the focus from pure output generation to reliable partnership.
Lalam: I think the improvement in measuring User Correction is vital because it directly correlates with how much human effort is required for success; if we can reduce that number, we’re seeing a direct improvement in the user experience.
Conclusion: Tom: So, wrapping up on "SWE-Together: Evaluating Coding Agents in Interactive User Sessions," the main implication is that evaluating coding agents needs to look at interaction quality and the human effort required for success, using User Correction as a key signal of capability.
Jane: Exactly. The paper shows that stronger models tend to require less corrective steering to reach comparable performance levels on tasks where instructions evolve during coding. It shifts the bar from just outputting syntactically correct code to being a more effective partner in a complex engineering conversation.
Lu: This work is important because it grounds the evaluation in real human behavior, moving us closer to models that can genuinely handle the iterative nature of software development tasks as we see them today. It gives us a better lens on what true agency looks like.
Meng: For deployment, this means we need systems that are reliable not just once, but across a sequence of interactions where requirements are fluid. We have to prioritize models that reduce that costly back-and-forth interaction time to make the whole process practical for real engineering work.
Lalam: The overall finding is that it’s about alignment; the goal is to design agents whose initial trajectory matches the user's evolving intent so we spend less time correcting them and more time building things. It really sets a new standard for what we expect from these AI assistants moving forward.
More episodes
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck
- 2407.14562-Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Bootstrapped Prolog-based Chain-of-Thought