GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes

arXiv:2606.28514 · cs.AI, cs.CL · Submitted 2026-06-26 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes".

Jane: The paper was written by Amit Parekh, Sabrina McCallum, Kareem Al-Hasan, Malvina Nikandrou and Alessandro Suglia from Heriot-Watt University and University of Edinburgh.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary of Findings: Tom: The authors provide a pretty detailed summary of how these models performed across their various experiments in GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes. The big finding is that, even in the simplest scenarios they tested, not a single model achieved success in real time.

Jane: It’s quite sobering to see that these state-of-the-art models fall far behind human players who can usually clear at least one module in a simple mission setup. The researchers are showing us where the current technology is lacking when it’s trying to cooperate effectively.

Lu: What I find particularly telling is the pattern of failure they observed; even if they can't finish a full mission, the models consistently struggle with basic tasks like solving individual modules in isolation, which points to deeper issues than just overall speed.

Meng: From an engineering viewpoint, it seems that the constraints of real-time asynchronous action are proving to be a huge bottleneck for current AI. The fact that they have no single model finish a full mission suggests that this isn't just a learning curve problem; it's structural.

Lalam: The paper also highlights how much more challenging the game is when you have to work with different models in an "other-play" setup compared to when using the same model for both roles. This tells us that collaboration requires a lot of flexibility from our AI.

Tom: So, Jane, it looks like this benchmark is designed to really expose those weaknesses in real time coordination rather than just testing if they can memorize solutions or relying on existing knowledge, right?

Suggested Improvements: Tom: This next part of the discussion focuses on what the authors think could improve both the models and our understanding of AI, Jane. It’s not just about how bad they are; at least, we see where they fail.

Jane: They have identified several low-level weaknesses in current models, such as struggling with state tracking across turns and having difficulty recovering from mistakes. This is a huge step because identifying these specific failure modes helps us know exactly what kind of training or architecture changes we need to make to fix them.

Lu: The way the paper suggests improving communication is particularly interesting; they are showing how much the models struggle with ambiguity and hallucination, especially when their partners provide information that's unclear. This highlights a crucial need for a robust mechanism to question assumptions in AI systems.

Meng: From an implementation side, I think the authors suggest we need better methods for the models to handle errors or "recover" from bad decisions instead of just giving up on a task after one strike. The current framework seems to penal us heavily when we make even a small mistake.

Lalam: The paper suggests that our AI needs better skills at understanding context and integrating information, not just by looking at the whole picture but by seeing how pieces fit together over time, which is very helpful for long-term cultural impact.

Tom: So, Jane, it seems like GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes is suggesting that we need models that are more than just good at following instructions; they need to be reliable partners who can adapt to a better human or AI team.

Implications: Tom: Let's move into the implications of this work, Jane, looking beyond what the paper found and thinking about what it means for future AI systems. The authors are presenting GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes as a way to measure things that current benchmarks simply ignore.

Jane: It’s a strong argument that this paper is forcing us to rethink what collaboration really looks like in the world. We often test single agents, but GPTNT demands genuine two-way interaction under pressure, which is essentially how real-world complex tasks work.

Lu: I see this as a major theoretical shift; it' moves us away from just asking AI to complete a task and toward asking AI to actually *partner* on a dynamic problem. This opens up massive possibilities for systems that require human-AI teamwork.

Meng: The implication for me is that it’ sets a new bar for performance in real-time agentic systems. If we are going to use AI alongside humans, we need systems that can function under this kind of time pressure and information asymmetry, not just static datasets.

Lalam: This work suggests that the future AI needs to be more than just computationally powerful; it needs to possess a shared understanding with its partners and exhibit reliable skills when we are working towards a goal together.

Tom: So, Jane, this really seems like the foundation for building systems that can actually coexist and cooperate with humans in complex ways.

Conclusion: Tom: As we wrap up our discussion of GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes, I think it’s clear that the paper has done a lot of work to define what real-time collaboration looks like for AI.

Jane: It's an incredibly valuable resource because it shows us the gap between current AI capabilities and what we need to achieve, providing a much clearer roadmap for improving our models than previous tests.

Tom: I want to hear final thoughts from Lu, Meng, and Lalam on this really, as we all wrap up the segment.

Lu: I think it’s fascinating how the creators have engineered this system to challenge is-ness of both open and closed source models that we're currently using for AI. It’s a true test of emergent intelligence in a dynamic setting.

Meng: We should be looking at this as an iterative process; the fact that you can keep generating new missions means the benchmark will evolve alongside our improvements, which is very practical for long-term testing.

Lalam: The advances in seeing how AI struggles with asynchronous coordination really push us toward a more nuanced and reliable form of human-AI interaction, which is a massive positive for the culture of how we use technology.

Tom: That's a great way to look at it, Lalam. Thank you all for this discussion on GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes today.

Heriot-Watt University · University of Edinburgh

cs.AI, cs.CL

Submitted: 2026-06-26

Updated: 2026-09-04

Comments: Accepted by TMLR on 02 Sept 2026. Project website and code at https://gptnt.github.io

Project page: https://gptnt.github.io

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 83/100

The gist: GPTNT is a novel benchmark designed to rigorously evaluate real-time, multimodal collaboration between artificial agents, addressing a significant gap in existing benchmarks that traditionally

Key concepts

Real-Time Collaboration
This refers to the ability of multiple AI agents to work together under immediate pressure and time constraints. The paper specifically tested this by requiring agents to cooperate in the complex, asynchronous environment of Keep Talking And Nobody Explodes.
Multimodal Agents
These are AI systems designed to process and utilize various types of information simultaneously. In this benchmark, the agents must handle multiple data sources—such as visual cues and textual instructions—to solve a complex problem together.
State Tracking and Recovery
This describes the AI's ability to maintain awareness of past actions and context across multiple turns. The paper highlighted that current models struggle with state tracking, meaning they often fail to recover or adjust after making an initial mistake.

Terminology

Summary

GPTNT is a novel benchmark designed to rigorously evaluate real-time, multimodal collaboration between artificial agents, addressing a significant gap in existing benchmarks that traditionally studied time pressure, information asymmetry, and imperfect communication in isolation. By adapting the cooperative video game Keep Talking and Nobody Explodes (ktane), the framework creates an ecologically valid environment where two agents must coordinate under strict constraints to solve procedurally generated bomb puzzles. The benchmark demonstrates that current state-of-the-art models fail to meet the demands of this real-world collaborative setting, providing a necessary tool for holistic evaluation of integrated MLLM skills.

How it works

The gptnt framework requires two distinct agent roles: the Defuser and the Expert. The Defuser has visual access to the bomb but lacks the instruction manual, while the Expert possesses full access to the manual but cannot see or manipulate any part of the bomb. This setup ensures that neither agent can succeed alone, making effective, efficient communication necessary for task completion. The benchmark utilizes a procedural generation system where each mission is defined by a set of constraints, including specific module types and time limits.

The real-time challenge

Unlike many existing benchmarks that rely on turn-taking protocols, gptnt operates in an asynchronous real-time mode where every inference counts against the clock. The bomb's countdown timer advances continuously while the models are generating their responses, meaning latency is a critical factor in the performance. This departs significantly from traditional agentic benchmarks, forcing agents to act asynchronously and communicate under genuine time pressure.

Evaluation and failure modes

We tested five state-of-the-art MLLMs (e.g., GPT-5.2, Claude Sonnet 4.6) on ten multi-module missions in the full setting. The results show that not one of the closed- and open-source models we test defuses a single bomb in real time, a bar human players can clear easily. Through systematic analysis, we identify critical weaknesses where models break down, including:

  • State tracking across multiple turns.

  • Efficient acting within the time budget.

  • Handling ambiguity in partner communication.

  • Error recovery from mistakes.

Diagnostic tools and contamination

To isolate genuine collaborative competence from simple memorization, gptnt includes specific diagnostic tools. We conducted single-agent evaluations where models had full access to both the bomb and the manual, confirming that communication is a source of error in the collaborative setting. Furthermore, we removed access to the instruction manual entirely to expose instances where models rely on pre-existing parametric knowledge of ktane's solutions, thereby surfacing reliance on memorised solutions rather than deriving answers in real time.

Improvements for AI systems

As an expert AI researcher, I have analyzed the provided paper on GPTNT (Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes). This work identifies critical failures in current state-of-the-art multimodal large language models (MLLMs) when they are tasked with real-time, asymmetric collaboration.

The following improvements address these systemic weaknesses, leading to a significantly more robust and effective AI system for complex collaborative tasks.


Based on the findings of GPTNT, we must move beyond simple single-turn task execution and implement structural changes that enforce temporal consistency and cognitive rigor in multi-agent settings.

Current agentic benchmarks often assume synchronous behavior (turn-taking). This creates a false sense of capability.

  • Improvement: The system must natively treat inference time as a first-class variable, not merely an artifact to be controlled for. This means modeling the bomb clock ticking down in real-time while the agent is deliberating or generating tokens (Asynchronous Mode).

  • Goal: Ensure agents can operate effectively under genuine temporal pressure, preventing lazy or overly verbose behavior from masking poor decision-making ability.

The paper shows models struggle to maintain state across turns (losing track of cumulative progress).

  • Improvement: Implement a recursive working memory mechanism that explicitly tracks all solved modules, the current strike count, and the precise time elapsed since every module's internal state change. This must be integrated into the ReAct framework’s thoughts block, ensuring that every action is grounded in a verifiable history.

  • Goal: Eliminate errors where agents fail to recognize when an action has successfully completed or when a previous mistake has been compounded by the passage of time.

Models often hallucinate visual information or misinterpret instructions, especially regarding visual cues like colors or markers.

  • Improvement: Enforce a dual-grounding requirement:

  • Set-of-Marks (SoM) as Primary: Use segmentation masks to reduce the action space to reliable, labeled interaction points.

  • Coordinate Verification: Implement a spatial check where any interaction must be verified against the precise pixel location, preventing models from selecting elements that are visually similar but spatially incorrect.

  • Goal: Ensure the AI system can reliably distinguish between visual ambiguity and accurate spatial positioning, making it less susceptible to superficial pattern matching.

Communication is often ad-hoc or relies on interpretation rather than verifiable data exchange.

  • Improvement: Implement a protocol for atomic messaging where all messages sent by the Expert are delivered in a single, non-summarized block to the Defuser, regardless of how long the Expert's generation takes. Furthermore, enforce clarification triggers: if an input (from visual or textual data) is ambiguous or contradicts known state (e.g., a serial number that doesn't have a vowel), the system must automatically generate a clarification query rather than guessing.

  • Goal: Prevent information loss and reduce the knowing-doing gap by forcing agents to address uncertainty rather than proceeding under false assumptions.

The paper highlights that models can leverage pre-training knowledge (memorizing the manual) without true reasoning.

  • Improvement: Integrate targeted, low-stakes offline VQA probes into the continuous evaluation loop. These probes test core competencies (like procedural logic or counting) independent of the full multi-turn game, ensuring that a successful single-step solution is not merely a form of parametric recall but requires demonstrable reasoning.

  • Goal: Force the AI to perform genuine, grounded reasoning rather than relying on memorized patterns derived from its training corpus.

A system incorporating these improvements would transform from a brittle, single-task solver into a robust, high-performance collaborative agent capable of:

  1. Sustained Collaborative Performance: The system can successfully navigate complex, multi-stage puzzles under real time pressure, maintaining consistency across multiple steps without forgetting previous solutions or mistakenly overthinking the task.

  2. Autonomous Error Recovery: When a model detects a strike (a failure state), it will not simply continue in the same flawed pattern. It will leverage its improved state tracking and clarification protocols to identify the root cause, hypothesize alternative strategies, and attempt recovery.

  3. High-Fidelity Role Execution: The system can reliably distinguish between its role as an Expert (systematically querying for necessary details) and a Defuser (systematically acting on visual cues), ensuring that its communication is always targeted, concise, and actionable.

  4. Robust Handling of Ambiguity: Instead of hallucinating or guessing when encountering unclear data (e.g., a blurry serial number or an ambiguous symbol), it will automatically pause and request clarification, maintaining a high standard of communicative integrity.

  5. Scalable Generalization: The system is prepared for environments that are not explicitly defined in its training data, as its core capabilities are grounded in the process of collaborative reasoning rather than memorized solutions to a specific set of puzzles.

Abstract

Multimodal models are increasingly deployed to solve tasks collaboratively with humans or other artificial agents. While existing benchmarks show that they possess the fundamental capabilities, the various conditions that coincide when collaborating---time pressure, information asymmetry, and imperfect communication---have traditionally been studied in isolation. To address this gap, we introduce GPTNT, a benchmark built on the cooperative video game Keep Talking and Nobody Explodes, in which two agents must coordinate to defuse procedurally generated bomb puzzles against a live countdown. One agent has access to the bomb but not the instructions for defusing it; the other holds the instructions but cannot see or manipulate the bomb. Neither agent can succeed alone: the task requires contributions from both, and is solvable only through effective, efficient communication. We remove turn-taking proxies or simplifications, instead requiring agents to act asynchronously and communicate in real time. GPTNT is designed to expose how models collaborate versus how they perform alone: the instruction manual, the partner, or both, can optionally be withheld to surface what a model has memorised versus what it derives in the moment. We demonstrate that GPTNT poses a considerable challenge to the state-of-the-art: not one of the closed- and open-source models we test defuses a single bomb in real time, a bar that human players clear. In a range of controlled experiments, we explore where capabilities break down, identifying critical weaknesses in state tracking, efficient acting within the time budget, handling ambiguity, and error recovery. Since it runs on the real game, GPTNT benefits from procedural generation and inherits a living modding community: as models improve, the benchmark can be evolved to remain challenging, rather than being solved once and retired.

Sources

Related papers