SWE-Together: Evaluating Coding Agents in Interactive User Sessions
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "SWE-Together: Evaluating Coding Agents in Interactive User Sessions".
Jane: Most coding-agent benchmarks are static, but real coding assistance is interactive, requiring users to clarify goals and correct mistakes over multiple turns.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, let's talk about those authors and what this paper is actually titled. The title of "SWE-Together: Evaluating Coding Agents in Interactive User Sessions" clearly signals the main focus here, which is on the interactive aspect of coding agent performance.
Jane: It’s quite direct, Tom. It tells us immediately that the research isn't about just seeing if an AI can write a piece of code, but how it behaves when things change during the process. That distinction between static and interactive is really important for understanding current capabilities.
Lu: The authors are clearly trying to bridge that gap between what we see in controlled, fixed benchmarks and what happens in real-world software engineering environments where users constantly pivot their needs. It's about capturing that dynamic flow of requirements.
Meng: From an engineering standpoint, the implication is that we need evaluation setups that simulate the actual workflow—where you have to pause and ask questions before moving on. That means testing isn't just running a script; it’s simulating a conversation with a junior developer who keeps changing their mind.
Lalam: I see it as a way to evaluate if an agent is truly becoming a collaborator, not just an autocomplete tool that spits out code based on the first prompt. If they can handle trajectory-conditioned replays, that suggests deeper understanding of the overall engineering goal.
The paper's summary: Tom: So, summarizing what SWE-Together does, it takes those raw, multi-turn coding agent sessions and turns them into reproducible software engineering tasks. They use a three-part pipeline to do this transformation before they even start the actual evaluation process.
Jane: That pipeline is where the heavy lifting happens; first, they filter and normalize those raw sessions to ensure we’re only looking at viable interactions, then they build a sandbox environment with pinned environments and executable checks. It’s a lot of preprocessing work to make sure the tasks are sound.
Lu: The second big piece is this user simulator that replays the original user's intent in a trajectory-conditioned manner, meaning it only steps in when the agent actually needs help based on what happened before, not just on a pre-set schedule. That’s sophisticated simulation design.
Meng: The third part is how they evaluate it: they separate measuring the final code quality from measuring how well the simulator behaves—things like "User Correction" and "Intent Coverage." It shows that correctness isn't the only thing we should care about when we talk about agent performance in real scenarios.
Lalam: That measurement of User Correction is really telling; it suggests that a model doesn't need constant steering to succeed, which points toward a more intuitive understanding of the user’s evolving intent. It’s about achieving alignment naturally over time.
The paper's improvements: Tom: The paper highlights several specific improvements in their methodology, primarily focusing on how they measure agent behavior through diagnostics like Intent Coverage and User Correction. They are looking at how well the simulator matches the original intents and how often users actually need to correct the agent.
Jane: They also introduced a unique way to assess correctness using an agentic rubric judge that gets updated based on running phases, ensuring that even the scoring mechanism is calibrated for fairness across different agents being tested. That makes it more rigorous than just looking at one final output score.
Lu: The paper suggests that we should focus on models that minimize corrective steering because they demonstrate a deeper grasp of the task trajectory. It’s not just about getting the right code; it’s about getting there efficiently with minimal back-and-forth.
Meng: From a practical standpoint, this means we can start optimizing for models that are robust enough to handle ambiguity and provide relevant feedback proactively rather than waiting for a major error. That shifts the focus from pure output generation to reliable partnership.
Lalam: I think the improvement in measuring User Correction is vital because it directly correlates with how much human effort is required for success; if we can reduce that number, we’re seeing a direct improvement in the user experience.
Conclusion: Tom: So, wrapping up on "SWE-Together: Evaluating Coding Agents in Interactive User Sessions," the main implication is that evaluating coding agents needs to look at interaction quality and the human effort required for success, using User Correction as a key signal of capability.
Jane: Exactly. The paper shows that stronger models tend to require less corrective steering to reach comparable performance levels on tasks where instructions evolve during coding. It shifts the bar from just outputting syntactically correct code to being a more effective partner in a complex engineering conversation.
Lu: This work is important because it grounds the evaluation in real human behavior, moving us closer to models that can genuinely handle the iterative nature of software development tasks as we see them today. It gives us a better lens on what true agency looks like.
Meng: For deployment, this means we need systems that are reliable not just once, but across a sequence of interactions where requirements are fluid. We have to prioritize models that reduce that costly back-and-forth interaction time to make the whole process practical for real engineering work.
Lalam: The overall finding is that it’s about alignment; the goal is to design agents whose initial trajectory matches the user's evolving intent so we spend less time correcting them and more time building things. It really sets a new standard for what we expect from these AI assistants moving forward.
Yifan Wu, Zhuokai Zhao, Songlin Li, Ho Hin Lee, Jiacheng Zhu, Shirley Wu, Tianhe Yu, Serena Li, Lizhu Zhang, Xiangjun Fan
cs.AI, cs.CL, cs.SE
Submitted: 2026-06-29
Updated: 2026-06-29
Code: https://github.com/Togetherbench/SWE-Together
Project page: https://togetherbench.com
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: Most coding-agent benchmarks are static, but real coding assistance is interactive, requiring users to clarify goals and correct mistakes over multiple turns.
Key concepts
- SWE-Together
- A benchmark built from 11,260 recorded user–agent coding sessions. It tests agents by measuring both the final correctness of their code and the number of corrective feedback turns needed during their interaction.
- Task Construction Pipeline
- A three-step process that transforms raw coding sessions into reproducible tasks. This involves filtering raw data, checking for local reproducibility, and using an LLM judge to confirm if a session can be turned into a verifiable task in an isolated sandbox.
- User Correction
- A diagnostic metric measuring how much the user simulator had to intervene. It tests the hypothesis that stronger models require less steering; lower correction scores suggest better models need less corrective guidance to succeed.
Terminology
Summary
Most coding-agent benchmarks are static, but real coding assistance is interactive, requiring users to clarify goals and correct mistakes over multiple turns. This paper introduces SWE-Together, a multi-turn benchmark reconstructed from 11,260 recorded user–agent coding sessions to evaluate agents as collaborators by measuring both final correctness and the number of corrective feedback turns required during interaction.
How it works
SWE-Together transforms raw multi-turn coding agent sessions into reproducible interactive software engineering tasks through a three-component pipeline. First, the task construction pipeline filters and normalizes raw sessions, screens for local reproducibility, and converts viable sessions into sandboxed repository-level tasks with pinned environments and executable checks. This stage involves:
-
A deterministic rule-based collector that filters raw upstream sessions based on criteria like containing multiple genuine user messages, concrete agent actions or code edits, and sufficient repository signal to identify the working repository.
-
An LLM judge that determines whether the coding work can be reconstructed as a reproducible and verifiable task.
-
A sandbox orchestrator that runs task construction to produce a complete task directory inside an isolated sandbox, where a task-generation agent performs a stricter repository-grounded screen, writes tests, and constructs user-simulation prompts.
How it works (Continued)
Second, the user simulator replays the original user’s intent in a trajectory-conditioned manner,
intervening only when conditions derived from the recorded session are satisfied. The simulator's action space includes no-op, question, redirect, new-requirement, and check-external. It follows two principles: interventions are trajectory-conditioned rather than scheduled
and anchored to the original session rather than generic.
This ensures that feedback timing is adapted to each agent's evolving trajectory while preserving the original user’s objectives and intervention order.
How it works (Continued)
Third, the evaluation framework separately measures correctness of the agent’s final repository state and user-simulator behavior. Task correctness is assessed using a combination of deterministic verifiers with an agentic rubric judge.
The rubric is fixed offline across agents to ensure comparability: Phase 1 runs once per task to derive a weighted task rubric, and Phase 2 applies that same rubric to each candidate repository state. The final score is derived mechanically: score = round X / w g I[g met], 2,
where weights are normalized.
How it works (Continued)
User-simulator behavior is characterized through two diagnostics: Intent Coverage and User Correction. Intent Coverage measures simulator fidelity by decomposing original session trajectories into atomic intents and matching the simulator’s messages against these intents, deriving IntentCoverage = round(0.70 Irecall + 0.30 Iprecision, 2).
User Correction tests the hypothesis that stronger models require less intervention; it is computed as UserCorrection = Ncorrection + 0.2 Nnudge,
where explicit corrections receive full weight and nudges receive a smaller weight.
How it works (Continued)
The evaluation reports four equally task-weighted correctness metrics: pass@1, SSR, pass2, and MeanJudge. The paper finds that User Correction is strongly inversely correlated with performance,
suggesting that stronger models require less corrective intervention to reach the same outcome.
Furthermore, Intent Coverage remains broadly stable across coding agents,
indicating cross-agent comparability.
How it works (Continued)
The experiments evaluate seven frontier models on the full 109-task benchmark using the opencode harness with k = 2 replicates per task. The results show that Claude Opus 4.8 achieves the strongest overall performance, leading in pass@1 (63%), SSR (59%), pass2 (52%), and mean judge score (0.801),
while requiring the least corrective steering with a mean User Correction of 1.38.
The simulator quality study found that human annotators exhibited no statistically reliable ability to distinguish simulated users from real users,
achieving a Turing pass rate of 46%.
How it works (Continued)
The paper distinguishes SWE-Together from prior work by combining repository-level agent-environment interaction, interactive user-correction replay, and provenance from real recorded user–agent coding sessions. This design choice is supported by the finding that simulators grounded in real human behavior yield substantially better downstream collaborative assistants than role-playing prompts.
The limitation noted is that the simulator cannot interrupt the coding agent during its turns or rely on visual information from the interface.
How it works (Continued)
The conclusion emphasizes that evaluating coding agents requires accounting for interaction quality and the human effort needed for success,
showing that User Correction is the one interaction signal that tracks capability.
SWE-Together aims to evaluate agents in a way that more closely reflects the actual user experience.
The overall finding is that stronger models need less steering to reach comparable performance.
Improvements for AI systems
Here are specific improvements to AI systems based on the SWE-Together benchmark and methodology described in the paper, along with what these improved systems can achieve:
The primary improvement is shifting from evaluating static, one-shot code generation to evaluating agents as active collaborators within complex, multi-turn engineering workflows.
A system that utilizes SWE-Together (or a similar interactive replay framework) can be trained or fine-tuned to exhibit superior behavior in environments requiring iterative refinement and clarification.
The improved AI system can be specifically engineered to reduce the number of required user interventions (User Correction) needed to achieve a high final success rate, indicating it is more capable of handling complex, evolving specifications autonomously.
By incorporating the metrics from the evaluation protocol, an AI system can be optimized not just for correctness but also for Intent Coverage.
This means the system will be trained to ensure that every underlying user goal and constraint introduced across multiple turns is accurately reflected in its final output, reducing drift
or omission of requirements.
The system can be designed to minimize unnecessary corrective feedback (User Correction) while maintaining high performance on task correctness. This suggests an improved user experience where the agent's initial understanding and trajectory are more aligned with the user's evolving intent, leading to fewer frustrating back-and-forth corrections.
An AI system can be evaluated for its Efficiency
(Wall-clock time and output tokens per task). The goal is to develop models that achieve high correctness scores with lower computational cost, making them more practical for real-world deployment where latency and resource consumption matter.
The system can be fine-tuned using the feedback loop derived from the User Simulator. This allows the agent to learn a policy of when to proactively ask clarifying questions or provide updates based on its current progress, mimicking a highly effective human collaborator rather than just a reactive code generator.
A system can be developed with robust verification pipelines that use deterministic verifiers alongside agentic rubrics (as described in Section 2.3.1). This ensures that the final code not only looks correct but is also executable and satisfies the specific behavioral goals of the original session, preventing false negatives
from superficially correct but functionally flawed solutions.
The system can be benchmarked against a baseline using metrics like Pass@2 (joint success across replicates), which heavily penalizes inconsistency between different agent runs on the same task. This drives development toward models with high reliability and stable performance rather than those that succeed only sporadically.
Abstract
Most coding-agent benchmarks are static: an agent receives a complete task description up front and is judged only by its final code. Real coding assistance is interactive, with users clarifying goals, adding constraints, and correcting mistakes over multiple turns. We introduce SWE-Together, a multi-turn benchmark reconstructed from real user-agent coding sessions. To make real interactions verifiable, we curate 109 repository-level tasks from 11,260 recorded sessions, selecting sessions with recoverable repository states, clear user goals, and observable outcomes. To replay these interactions across agents, we build a reactive LLM-based user simulator that preserves the original users' intents and provides feedback when the coding agent's progress requires it. To evaluate agents as collaborators, we measure both final repository correctness and the number of corrective feedback turns required during the interaction. Experiments with frontier coding agents show that stronger agents generally achieve higher final success rates while requiring fewer interventions, suggesting an improved user experience.
Sources
- Program Synthesis with Large Language Models
- SWE-chat: Coding Agent Interactions From Real Users in the Wild
- Evaluating Large Language Models Trained on Code
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection