GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes
summary
The gist
GPTNT is a novel benchmark designed to rigorously evaluate real-time, multimodal collaboration between artificial agents, addressing a significant gap in existing benchmarks that traditionally
In short
The episode discusses the GPTNT paper, which benchmarks real-time collaboration between multimodal AI agents using the game Keep Talking And Nobody Explodes. The hosts conclude that current state-of-the-art models fail to achieve success in simple, time-pressured scenarios, suggesting a need for AI systems capable of genuine two-way partnership rather than just following instructions.
Key concepts
- Real-Time Collaboration
- This refers to the ability of multiple AI agents to work together under immediate pressure and time constraints. The paper specifically tested this by requiring agents to cooperate in the complex, asynchronous environment of Keep Talking And Nobody Explodes.
- Multimodal Agents
- These are AI systems designed to process and utilize various types of information simultaneously. In this benchmark, the agents must handle multiple data sources—such as visual cues and textual instructions—to solve a complex problem together.
- State Tracking and Recovery
- This describes the AI's ability to maintain awareness of past actions and context across multiple turns. The paper highlighted that current models struggle with state tracking, meaning they often fail to recover or adjust after making an initial mistake.
Terminology used across episodes
This episode discusses
- GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes · Paper Radio
- ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems
- SpurLens: Automatic Detection of Spurious Cues in Multimodal LLMs
- AvalonBench: Evaluating LLMs Playing the Game of Avalon
- VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
- Dota 2 with Large Scale Deep Reinforcement Learning
- BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- Needle In A Multimodal Haystack
- Enhance Reasoning for Large Language Models in the Game Werewolf
The paper
GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes · Read on arXiv
Heriot-Watt University · University of Edinburgh
Multimodal models are increasingly deployed to solve tasks collaboratively with humans or other artificial agents. While existing benchmarks show that they possess the fundamental capabilities, the various conditions that coincide when collaborating---time pressure, information asymmetry, and imperfect communication---have traditionally been studied in isolation. To address this gap, we introduce GPTNT, a benchmark built on the cooperative video game Keep Talking and Nobody Explodes, in which two agents must coordinate to defuse procedurally generated bomb puzzles against a live countdown. One agent has access to the bomb but not the instructions for defusing it; the other holds the instructions but cannot see or manipulate the bomb. Neither agent can succeed alone: the task requires contributions from both, and is solvable only through effective, efficient communication. We remove turn-taking proxies or simplifications, instead requiring agents to act asynchronously and communicate in real time. GPTNT is designed to expose how models collaborate versus how they perform alone: the instruction manual, the partner, or both, can optionally be withheld to surface what a model has memorised versus what it derives in the moment. We demonstrate that GPTNT poses a considerable challenge to the state-of-the-art: not one of the closed- and open-source models we test defuses a single bomb in real time, a bar that human players clear. In a range of controlled experiments, we explore where capabilities break down, identifying critical weaknesses in state tracking, efficient acting within the time budget, handling ambiguity, and error recovery. Since it runs on the real game, GPTNT benefits from procedural generation and inherits a living modding community: as models improve, the benchmark can be evolved to remain challenging, rather than being solved once and retired.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes".
Jane: The paper was written by Amit Parekh, Sabrina McCallum, Kareem Al-Hasan, Malvina Nikandrou and Alessandro Suglia from Heriot-Watt University and University of Edinburgh.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Summary of Findings: Tom: The authors provide a pretty detailed summary of how these models performed across their various experiments in GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes. The big finding is that, even in the simplest scenarios they tested, not a single model achieved success in real time.
Jane: It’s quite sobering to see that these state-of-the-art models fall far behind human players who can usually clear at least one module in a simple mission setup. The researchers are showing us where the current technology is lacking when it’s trying to cooperate effectively.
Lu: What I find particularly telling is the pattern of failure they observed; even if they can't finish a full mission, the models consistently struggle with basic tasks like solving individual modules in isolation, which points to deeper issues than just overall speed.
Meng: From an engineering viewpoint, it seems that the constraints of real-time asynchronous action are proving to be a huge bottleneck for current AI. The fact that they have no single model finish a full mission suggests that this isn't just a learning curve problem; it's structural.
Lalam: The paper also highlights how much more challenging the game is when you have to work with different models in an "other-play" setup compared to when using the same model for both roles. This tells us that collaboration requires a lot of flexibility from our AI.
Tom: So, Jane, it looks like this benchmark is designed to really expose those weaknesses in real time coordination rather than just testing if they can memorize solutions or relying on existing knowledge, right?
Suggested Improvements: Tom: This next part of the discussion focuses on what the authors think could improve both the models and our understanding of AI, Jane. It’s not just about how bad they are; at least, we see where they fail.
Jane: They have identified several low-level weaknesses in current models, such as struggling with state tracking across turns and having difficulty recovering from mistakes. This is a huge step because identifying these specific failure modes helps us know exactly what kind of training or architecture changes we need to make to fix them.
Lu: The way the paper suggests improving communication is particularly interesting; they are showing how much the models struggle with ambiguity and hallucination, especially when their partners provide information that's unclear. This highlights a crucial need for a robust mechanism to question assumptions in AI systems.
Meng: From an implementation side, I think the authors suggest we need better methods for the models to handle errors or "recover" from bad decisions instead of just giving up on a task after one strike. The current framework seems to penal us heavily when we make even a small mistake.
Lalam: The paper suggests that our AI needs better skills at understanding context and integrating information, not just by looking at the whole picture but by seeing how pieces fit together over time, which is very helpful for long-term cultural impact.
Tom: So, Jane, it seems like GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes is suggesting that we need models that are more than just good at following instructions; they need to be reliable partners who can adapt to a better human or AI team.
Implications: Tom: Let's move into the implications of this work, Jane, looking beyond what the paper found and thinking about what it means for future AI systems. The authors are presenting GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes as a way to measure things that current benchmarks simply ignore.
Jane: It’s a strong argument that this paper is forcing us to rethink what collaboration really looks like in the world. We often test single agents, but GPTNT demands genuine two-way interaction under pressure, which is essentially how real-world complex tasks work.
Lu: I see this as a major theoretical shift; it' moves us away from just asking AI to complete a task and toward asking AI to actually *partner* on a dynamic problem. This opens up massive possibilities for systems that require human-AI teamwork.
Meng: The implication for me is that it’ sets a new bar for performance in real-time agentic systems. If we are going to use AI alongside humans, we need systems that can function under this kind of time pressure and information asymmetry, not just static datasets.
Lalam: This work suggests that the future AI needs to be more than just computationally powerful; it needs to possess a shared understanding with its partners and exhibit reliable skills when we are working towards a goal together.
Tom: So, Jane, this really seems like the foundation for building systems that can actually coexist and cooperate with humans in complex ways.
Conclusion: Tom: As we wrap up our discussion of GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes, I think it’s clear that the paper has done a lot of work to define what real-time collaboration looks like for AI.
Jane: It's an incredibly valuable resource because it shows us the gap between current AI capabilities and what we need to achieve, providing a much clearer roadmap for improving our models than previous tests.
Tom: I want to hear final thoughts from Lu, Meng, and Lalam on this really, as we all wrap up the segment.
Lu: I think it’s fascinating how the creators have engineered this system to challenge is-ness of both open and closed source models that we're currently using for AI. It’s a true test of emergent intelligence in a dynamic setting.
Meng: We should be looking at this as an iterative process; the fact that you can keep generating new missions means the benchmark will evolve alongside our improvements, which is very practical for long-term testing.
Lalam: The advances in seeing how AI struggles with asynchronous coordination really push us toward a more nuanced and reliable form of human-AI interaction, which is a massive positive for the culture of how we use technology.
Tom: That's a great way to look at it, Lalam. Thank you all for this discussion on GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes today.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language