Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration".
Jane: The paper was written by Yijia Shao, Vinay Samuel, Yucheng Jiang, John Yang and Diyi Yang from Stanford University and Carnegie Mellon University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, folks! We've got a paper that's been making waves in the AI world, and the title alone tells you a lot: "Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration." Jane, when you first saw that title, what jumped out at you?
Jane: Honestly, Tom, it was the word "Collaborative." So much of the AI research we talk about is about autonomous agents doing things on their own, like a robot vacuum that just goes. This paper is saying, "No, wait, what if the robot vacuum and the human work together to clean the house?" It's a shift in the whole philosophy.
Tom: And that's a big deal, right? It's from a team at Stanford and Carnegie Mellon, led by Yijia Shao and Diyi Yang. They're basically building the gymnasium where these human-AI teams can train and be tested.
Lu: From my perspective, the title is a promise. It's promising a structured way to study a messy, human problem. We've all had frustrating experiences with chatbots that just don't get it. This framework is an attempt to move past that, to give us a scientific way to measure and improve that interaction.
Meng: As an engineer, I see the title and I immediately think about the infrastructure. "Gym" is a great metaphor because it implies a controlled environment with rules, tools, and a way to keep score. You need that to build reliable systems. You can't just throw a model into the wild and hope for the best.
Jane: Exactly, Meng. And the title also hints at the "dual control" aspect, which is the real heart of it. It's not just the agent doing stuff and the human watching. It's a shared workspace where both can act. Like, you're both editing the same document at the same time.
Tom: So it's not a back-and-forth conversation like we're having now. It's more like two people working on a project together on a whiteboard.
Jane: Precisely. And that's what makes this paper so exciting to me. It's not just about making agents smarter; it's about making them better teammates. And the implications for the world, for how we'll work with these systems, are huge.
Lu: The implications are huge, but the challenges are equally huge. The paper's title suggests a solution, but it also implies a question: How do you even begin to evaluate something as subjective as collaboration? That's what I'm eager to dig into.
Tom: You're reading my mind, Lu. Let's not just stare at the title. Let's get into the meat of the paper and see how they actually built this gym and what they found inside.
Summary: Jane: So, we've got the title, which is all about teamwork. Now let's talk about what the paper actually does. "Collaborative Gym" isn't just a concept; it's a real, open-source framework with code and data. And the core idea is to move beyond turn-taking, you know, like a chatbot where you say something, it replies, you say something else.
Tom: Right, it's like ping-pong. But this paper is saying, "Forget ping-pong. Let's play a game of basketball where everyone is on the court at the same time." They call it "non-turn-taking interaction."
Lu: And that's a crucial distinction. In a real team, you don't wait for your teammate to finish their entire task before you start yours. You might be writing a report while your teammate is searching for data. This framework allows for that kind of parallel, asynchronous work.
Meng: The paper calls this "dual control" in a shared workspace. The human and the agent both have their hands on the wheel. The agent can update a document, and the human can jump in and edit the same document or send a message without waiting for the agent to finish.
Jane: Exactly. And to make that work, they had to build a whole new coordination protocol. It's not just about the actions, but about the notifications. The agent needs to know when the human has changed something, and the human needs to see what the agent is doing in real time.
Tom: And they built this on three different tasks to test it out. We're talking about planning a trip, writing a "Related Work" section for a paper, and analyzing a bunch of tabular data. Each one has its own challenges, right?
Jane: Oh, for sure. Travel planning is all about those hidden preferences. You know, the user might have a secret budget or a specific vibe they're going for. Writing related work is about domain expertise. And tabular analysis is about combining the agent's coding skills with the human's understanding of the data.
Lu: The most striking result for me is that the best collaborative agents consistently beat their fully autonomous counterparts. In the real-world tests, they had win rates of eighty-six percent in travel planning and seventy-four percent in tabular analysis. That's a massive vote for the power of collaboration.
Meng: It is, but it's not a free lunch. The paper is very honest about the failures too. They found that communication and situational awareness are huge bottlenecks. The agents often fail to keep the human in the loop, or they lose track of the context of the conversation.
Jane: So, the summary is that collaboration is demonstrably better, but our current AI models are still not very good at it. They need the right framework to even have a chance, and even then, they stumble.
Tom: So we've got the "what" and the "why". But the paper doesn't just stop at showing it's a good idea. It actually suggests some specific ways to make these agents better. Let's talk about those improvements next.
Improvements: Tom: Alright, so we know that human-agent teams can beat solo agents, but the agents themselves have some serious flaws. The paper doesn't just point these out; it actually tests a specific improvement. What was that, Jane?
Jane: They call it the "Collaborative Agent with Situational Planning." The idea is to give the agent a moment to think about the bigger picture before it acts. Instead of just reacting to the latest notification, it's prompted to make a three-way decision: should I take a task action, send a message, or just wait?
Lu: And that "wait" option is so important. It's the agent acknowledging that the human might have something to say or do. It's a form of self-awareness, a way to avoid stepping on the human's toes. This simple addition made a huge difference in the balance of the collaboration.
Meng: From an engineering standpoint, it's a clever way to add a planning layer without changing the underlying model. You're just changing the prompt. And the results speak for themselves. This version of the agent significantly outperformed the baseline collaborative agent, especially in the simulated tests.
Jane: Right, and it also improved the "Initiative Entropy," which is their metric for how balanced the conversation is. In a good collaboration, both parties are taking initiative. The agent isn't just a passive tool, and the human isn't just giving orders. They're both contributing ideas and directions.
Tom: So, it's not just about making the agent smarter, but about making it a better listener and a more thoughtful teammate. But the paper also points out that the framework itself can be improved. What about the interaction paradigm? They compared it to a turn-taking version, right?
Jane: Yes, and this is where it gets really interesting. They took the exact same agent and just changed the interaction protocol. One version was turn-taking, and the other was the non-turn-taking "Co-Gym" style. The non-turn-taking version won seventy percent of the time in head-to-head comparisons.
Meng: The users in the study said it felt more natural and more like working with a real teammate. They could step in and guide the agent at any moment, which they really valued. It's a strong signal that the way we interact with these systems is just as important as the model's raw intelligence.
Lu: So the improvement isn't just a new algorithm. It's a new philosophy of interaction. We're moving from a master-servant dynamic to a peer-to-peer dynamic. And the paper shows that this philosophical shift has concrete, measurable benefits.
Tom: So we've got a framework, we've got a better agent design, and we've got a better interaction model. But all of this is built on the first page of the paper, which lays out the whole problem. Let's go back to the beginning and see how they frame this whole field.
First Page: Jane: We've been talking about the results and the improvements, but it all starts with the problem they identify on the very first page of "Collaborative Gym." They make a really sharp observation: most AI agent research is focused on full automation.
Tom: Right, the dream of just giving an agent a task and letting it run off and do everything by itself. But this paper argues that's not always the right goal. There are tons of real-world situations where you *need* a human in the loop because of their preferences, their expertise, or just because they want to stay in control.
Lu: Exactly. And they make a fantastic point about complementary expertise. The human might know what a good travel itinerary looks like, but the agent can search through thousands of options in seconds. The agent can write code, but the human knows what the data actually means. The best results come from combining those strengths.
Meng: And the first page also introduces the core components of their solution. They talk about the "collaboration-driven environment design," which is their way of saying the task environment itself has to support two people working on it at the same time. It's not just a single-player game anymore.
Jane: They also introduce the idea of the "notification protocol," which is the technical backbone for that non-turn-taking interaction we talked about. It's how the agent knows when the human has made a change, and vice-versa. It's the glue that holds the whole collaboration together.
Tom: So the first page is really about setting the stage. It's saying, "Look, we've been building these amazing single-player agents, but the real world is a multiplayer game. And we need a new kind of infrastructure to play it."
Lu: And they also preview their evaluation suite, which is so important. They're not just asking, "Did the agent finish the task?" They're asking, "How good was the final result?" and "How good was the process of getting there?" That second question is what's truly novel here.
Meng: The paper's first page is a manifesto for a new kind of AI. It's a call to move from building tools to building teammates. And the rest of the paper is a compelling argument that this is not only possible but necessary.
Jane: And that's the perfect setup for our conclusion. We've seen the problem, the framework, the results, and the improvements. Now let's wrap it all up and see what it means for the future.
Conclusion: Tom: Well, folks, we've had a fantastic time unpacking "Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration." Jane, if you had to give our listeners one final takeaway, what would it be?
Jane: I think it's that collaboration isn't just a nice-to-have feature for AI; it's a necessity for tackling complex, real-world problems. The paper proves that human-agent teams can outperform either one working alone, but it also shows that our current models are still pretty clumsy at it.
Lu: And that's the exciting part for me. The paper identifies the bottlenecks so clearly—communication and situational awareness—which gives the whole research community a roadmap. We know exactly where to focus our efforts to make the next generation of AI much better at working with us.
Meng: From my side, the fact that they released the entire framework as open-source is a game-changer. It means any researcher or developer can take this "gym" and start training their own agents. It lowers the barrier to entry for this entire field of study.
Jane: And it's not just for researchers. The implications for everyday life are huge. Imagine a travel agent that actually listens to your preferences, or a writing assistant that understands your style, or a data analyst that can explain its findings to you in plain English. That's the world this paper is pointing toward.
Tom: So, as we say goodbye to this paper, we're not just closing a chapter. We're opening a new one. The era of the autonomous agent might be giving way to the era of the collaborative agent. And that's a future we should all be excited about.
Jane: Absolutely, Tom. It's a future where AI doesn't replace us, but works *with* us. And with frameworks like Collaborative Gym leading the way, that future is looking a lot closer.
Tom: Thanks for joining us, everyone. We'll see you next time on the show. Goodbye, Collaborative Gym!
Yijia Shao, Vinay Samuel, Yucheng Jiang, John Yang, Diyi Yang
Stanford University · Carnegie Mellon University
cs.AI, cs.CL, cs.HC
Submitted: 2026-08-08
Updated: 2026-08-11
Comments: ICLR 2026
Code: https://github.com/SALT-NLP/collaborative-gym
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 65/100
The gist: "creating travel plans, writing Related Work sections, and analyzing tabular data." The paper evaluates "three types of agent architectures—fully autonomous agents that only interact with the
Key concepts
- Collaborative Gym
- This is an open-source framework built to allow humans and AI agents to train and be tested together. It moves beyond simple back-and-forth turn-taking, providing a structured environment where the agent and data can work asynchronously.
- Non-Turn-Taking Interaction
- This concept describes a shared workspace where both parties act simultaneously. Instead of waiting for one to finish, both the human and have hands on the wheel, allowing them to update documents or work on tasks in parallel.
- Collaborative Agent with Situational Planning
- This is a specific design improvement where the agent is prompted to consider the bigger picture before acting. It must decide whether to take an action, send a message, or wait, making it a more thoughtful and balanced teammate.
Terminology
Summary
Summary
The paper introduces Collaborative Gym (Co-Gym), described as the first framework for developing and evaluating agents that can engage in bidirectional communication with humans while interacting with task environments.
The authors, from Stanford University and Carnegie Mellon University, address a gap in LM agent research where the human role remains largely overlooked.
Co-Gym consists of three core components. First, collaboration-driven environment design
which imposes no constraints on agent implementation but instead defines an environment interface that supports dual control in a shared workspace.
Second, a protocol for non-turn-taking interaction
where humans and agents coordinate through two collaboration acts and a notification protocol for real-time change monitoring rather than enforced turn structures.
Third, an evaluation suite that assesses both outcomes and processes,
evaluating task delivery and performance while auditing initiative-taking patterns and human satisfaction during collaboration.
The framework is instantiated in both simulated conditions with a reliable LM-based user simulator and real conditions with an interactive web application featuring a chat panel and shared workspace.
Three representative tasks are supported: creating travel plans, writing Related Work sections, and analyzing tabular data.
The paper evaluates "three types of agent architectures—fully autonomous agents that only interact with the environment, collaborative agents, and collaborative agents enhanced with a situational planning module—powered by four LMs under both simulated and real conditions."
Key results from the simulated condition show that "collaborative agents have a lower delivery rate as human-agent teams sometimes fail to reach their goals within the step limit due to poor communication or coordination; but among the delivered cases, human-agent teams tend to achieve higher-quality outcomes. Specifically, the
Collaborative Agent with Situational Planning achieves the best Task Performance across all three tasks and has significant improvement over the Fully Autonomous Agent powered by the same LM."
In the real condition, results indicate "human-agent collaboration can be beneficial, with the best-performing collaborative agent achieving win rates of 86% in Travel Planning and 74% in Tabular Analysis compared to fully autonomous agents when evaluated by real users. The paper also notes that
users voluntarily allocate 21-32% of their actions to direct environment control, demonstrating that
users actively value and exploit this complementary interaction channel."
Error analysis reveals persistent limitations in current language models and agents, with communication and situational awareness failures observed in 65% and 40% of cases in the real condition, respectively.
The error distributions between real and simulated conditions have a Spearman rank correlation of 0.803 (p = 7.94 × 10−7),
indicating strong alignment.
An ablation study comparing Co-Gym with a turn-taking version found that Co-Gym achieves higher task performance (0.84 vs 0.81), greater overall satisfaction (3.95 vs 3.55, p = 0.012 in pairwise t-test), and a substantial preference advantage (70% win rate)
in Travel Planning, with a similar trend in Tabular Analysis.
The paper concludes that "human-agent teams achieve superior performance compared to fully autonomous agents, while simultaneously exposing critical challenges in communication, situational awareness, and planning that necessitate advancements in both underlying LMs and agent design. Co-Gym is
released under the permissive MIT license to
facilitate the study of collaborative agents and
advances the broader goal of creating AI systems that augment human capabilities rather than replace them."
Improvements for AI systems
Based on the paper Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration,
here are the specific improvements I can implement in AI systems and the resulting capabilities:
Improvement: Implement a coordination protocol with two collaboration acts (SendTeammateMessage and WaitTeammateContinue) and a notification system that allows asynchronous, parallel interaction between humans and agents.
What the improved AI system can do:
-
Process multiple user messages simultaneously without waiting for responses
-
Send proactive updates and questions without requiring prior human input
-
Monitor shared workspaces and notify all parties of changes in real-time
-
Continue working on tasks while the human reviews or edits parts of the output
Sources
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- OpenAI Gym
- WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
- MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
- CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation
- DiscoveryBench: Towards Data-Driven Discovery with Large Language Models
- Long-Term Planning and Situational Awareness in OpenAI Five
- PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action
- ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
- On the Perception of Difficulty: Differences between Humans and AI
- Cognitive Architectures for Language Agents
- AutoSurvey: Large Language Models Can Automatically Write Surveys
- ReAct: Synergizing Reasoning and Acting in Language Models
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection