High Volatility and Action Bias Distinguish LLMs from Humans in Group Coordination
summary
The gist
Humans exhibit remarkable abilities to coordinate in groups, and this study investigates whether large language models (LLMs) can demonstrate comparable adaptive coordination by comparing their
In short
This study compared human and Large Language Model (LLM) performance in a common-interest game called Group Binary Search (GBS). Humans improved their coordination over multiple games, while LLMs showed minimal learning and excessive switching. LLMs overreacted to feedback, maintained high volatility, and failed to stabilize group behavior.
Key concepts
- Group Binary Search (GBS) Task
- Participants in this game must collectively guess a target number by submitting numerical guesses without talking. They rely only on group feedback—like 'too high' or 'too low'—to adjust their guesses iteratively until the correct total is reached.
- Excessive Switching
- LLMs frequently change their strategy or action in every round of the game. Humans, especially in larger groups, tend to reduce this switching as the group gets closer to the target. This high rate of change in LLMs impairs their ability to coordinate effectively.
- Overreactivity
- LLM agents tend to respond too strongly or too quickly when they receive feedback from the group. This contrasts with humans, who modulate their adjustments more carefully. The paper found LLMs exhibited a mean overreaction score of -1.386, suggesting instability.
- Action Bias
- LLMs have a persistent tendency to change their output or take an action in every round, even when stability would be better for the group. This bias suggests LLMs are implicitly biased toward change as progress toward the goal, rather than maintaining a stable position.
Terminology used across episodes
This episode discusses
- High Volatility and Action Bias Distinguish LLMs from Humans in Group Coordination · Paper Radio
- The Llama 3 Herd of Models · Paper Radio
- LLMs Get Lost In Multi-Turn Conversation
- DeepSeek-V3 Technical Report
- HyperAgent: Generalist Software Engineering Agents to Solve Coding Tasks at Scale
- Multi-Agent Collaboration Mechanisms: A Survey of LLMs
The paper
High Volatility and Action Bias Distinguish LLMs from Humans in Group Coordination · Read on arXiv
Department of Computer Science, Indiana University Bloomington · Cognitive Science Program, Indiana University Bloomington · Department of Psychological and Brain Sciences, Indiana University Bloomington
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "High Volatility and Action Bias Distinguish LLMs from Humans in Group Coordination".
Tom: Humans exhibit remarkable abilities to coordinate in groups,
Jane: First, who's behind it and why it matters.
Paper summary: Tom: Well, Jane, we're diving into this paper today titled "High Volatility and Action Bias Distinguish LLMs from Humans in Group Coordination." The main idea here is to see if these large language models can actually coordinate their actions as effectively as humans do in a situation where they have to work together toward a common goal.
Jane: It sounds like the core question is whether the AI can adapt and stabilize its behavior over time when coordinating with others, which is something we observe in human groups. The authors are testing this using a specific type of game called Group Binary Search, where everyone has to guess a target number together without talking directly.
Lu: I think what really interests me about this is how they frame the comparison between LLMs and humans in this setting. They aren't just looking at whether the AI can reach the target, but *how* it does it, focusing on things like volatility and action bias.
Meng: From an engineering standpoint, I’m curious about how they set up this common-interest game; for us to build useful coordination systems, we need to understand the constraints of those interactions. So what is their central claim about the performance difference?
Lalam: I see a massive potential here for how we design AI systems that interact in complex environments; if we can understand these coordination failures, it helps us build models that are more robust and culturally aligned with human group dynamics.
Tom: So, to summarize what the paper claims about the findings, it seems the central point is that unlike humans who improve and settle their behavior across multiple games, LLMs often struggle because they show excessive switching and overreactivity.
Jane: Exactly. The paper finds that LLMs frequently fail to improve their performance from one game session to the next, which is a big deal when we think about learning from experience within a team setting.
Lu: They specifically point out that LLMs exhibited minimal cross-game learning, whereas human groups generally show improvement for both directional and numerical feedback across successive games. This suggests a fundamental difference in how they learn coordination strategies.
Meng: That lack of learning is concerning because it means an AI team wouldn't naturally get better at coordinating as the tasks become more complex or as the group size changes. How does that translate into practical application for, say, a logistics planning system?
Lalam: If we can see this volatility and switching, it tells us we need to engineer mechanisms that dampen excessive reactivity in our AI agents so they don't introduce instability when working with human operators or other AI systems.
Tom: And the paper drills down into *why* this happens, suggesting LLMs have a consistent bias toward changing their output in every round, which the authors call an action bias.
Paper summary: Jane: That action bias seems to be a major culprit because it means the AI keeps taking actions even when staying stable might be better for group convergence, leading to those oscillations around the target they mentioned.
Lu: The paper quantified this by showing that human groups tended to underreact but modulate their adjustments, while LLMs showed a mean slope of-one point three eight six across conditions, which the authors interpret as overreaction. That numerical difference is significant.
Meng: A consistently negative slope in the LLM adaptation data suggests that the model itself is struggling to find a stable policy rather than just making small, correct adjustments, which brings up practical implementation issues regarding reward shaping.
Lalam: This points toward a need for architectures that inherently prioritize stability and long-term goals over immediate reaction, which could fundamentally change how we structure decision-making in complex operational environments.
Tom: Moving on to the conclusion of this study, the authors highlight that even when LLMs are given rich numerical feedback, they don't show the same benefit as humans do when that feedback is directional.
Jane: It really emphasizes that simple adjustment isn't enough; it’s about how those adjustments are modulated and stabilized over time in a shared context.
Lu: The authors are concluding that the persistent volatility, excessive switching, and overreactivity we observed provide a behavioral diagnostic for understanding the coordination gap between LLMs and humans.
Meng: So, the implication is that simply making an LLM more capable in terms of reasoning won't automatically make it better at coordinating tasks like this without addressing these specific behavioral patterns.
Lalam: I think the biggest impact for me is realizing that improving coordination isn't just about feeding more data to a model; it requires building in mechanisms that encourage stability and thoughtful response modulation within the AI architecture itself.
Tom: So, to wrap up this discussion on "High Volatility and Action Bias Distinguish LLMs from Humans in Group Coordination," we see that current LLMs lack the adaptive learning pattern humans exhibit across games, instead showing persistent volatility and action bias.
Jane: It’s clear that while they can perform individual tasks well, their ability to coordinate as a group is currently limited by these behavioral tendencies toward overreaction and switching.
Lu: This work opens up interesting avenues for research into how we can program stability directly into the coordination layers of large models.
Meng: For practical deployment, this suggests that before deploying an AI team for collaborative tasks, we need rigorous testing specifically focused on these kinds of group coordination dynamics rather than just individual accuracy metrics.
Lalam: This research is valuable because it gives us a concrete behavioral fingerprint to work with when trying to align AI behavior with the principles of effective human teamwork.
Conclusion: Tom: So, we’ve seen how these LLMs struggle to learn and adapt in group settings during this discussion of "High Volatility and Action Bias Distinguish LLMs from Humans in Group Coordination."
Jane: That paper really zeroes in on the fact that human groups show a consistent pattern of improvement across games, whereas the AI tends to swing wildly between actions without settling.
Lu: From my side, I’m fascinated by how they quantified that overreaction—that mean slope of-one point three eight six is quite telling about the underlying model behavior we're seeing in these systems.
Meng: It means if we build a coordination system on this foundation, we have to explicitly design stabilizers because the default behavior is to keep changing things too much.
Lalam: I see this as a huge step toward making AI teams more reliable; understanding this volatility helps us build models that actually function cohesively rather than just reacting blindly.
Tom: Exactly! The authors are showing us that the difference isn't just about the complexity of the task, but about how agents manage their internal state and feedback over time.
Jane: It really boils down to this: humans learn to modulate their adjustments as they get closer to a solution, while these LLM groups just keep overreacting.
Lu: And that persistence of switching, even when convergence is near, suggests there’s a deep bias built into how the models process feedback loops in this environment.
Meng: For me, the practical implication is that we need to focus our engineering efforts on reducing that intrinsic bias in how the AI prioritizes staying stable over making a move.
Lalam: If we can tame that action bias, I think it could fundamentally change how complex AI systems are allowed to operate in any collaborative setting.
Tom: This paper is a must-read for anyone working on multi-agent systems because it gives us a concrete behavioral reason why coordination fails when we try to automate it without this specific tuning.
Jane: It’s about moving past the idea that more processing power automatically means better group behavior, which this study really challenges.
Lu: We need to look into how the authors suggest integrating feedback history directly into the model architecture to address these learning deficits they observed across games.
Meng: That suggests we might need a new layer of control specifically designed for dampening volatility, something that goes beyond standard reinforcement learning setups.
Lalam: Thinking about the future, this work could inspire entirely new ways to structure collaborative AI workflows where stability is a core design principle from the start.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck