High Volatility and Action Bias Distinguish LLMs from Humans in Group Coordination

arXiv:2604.02578 · cs.MA, cs.AI, cs.CL, cs.GT · Submitted 2026-04-02 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "High Volatility and Action Bias Distinguish LLMs from Humans in Group Coordination".

Tom: Humans exhibit remarkable abilities to coordinate in groups,

Jane: First, who's behind it and why it matters.

Paper summary: Tom: Well, Jane, we're diving into this paper today titled "High Volatility and Action Bias Distinguish LLMs from Humans in Group Coordination." The main idea here is to see if these large language models can actually coordinate their actions as effectively as humans do in a situation where they have to work together toward a common goal.

Jane: It sounds like the core question is whether the AI can adapt and stabilize its behavior over time when coordinating with others, which is something we observe in human groups. The authors are testing this using a specific type of game called Group Binary Search, where everyone has to guess a target number together without talking directly.

Lu: I think what really interests me about this is how they frame the comparison between LLMs and humans in this setting. They aren't just looking at whether the AI can reach the target, but *how* it does it, focusing on things like volatility and action bias.

Meng: From an engineering standpoint, I’m curious about how they set up this common-interest game; for us to build useful coordination systems, we need to understand the constraints of those interactions. So what is their central claim about the performance difference?

Lalam: I see a massive potential here for how we design AI systems that interact in complex environments; if we can understand these coordination failures, it helps us build models that are more robust and culturally aligned with human group dynamics.

Tom: So, to summarize what the paper claims about the findings, it seems the central point is that unlike humans who improve and settle their behavior across multiple games, LLMs often struggle because they show excessive switching and overreactivity.

Jane: Exactly. The paper finds that LLMs frequently fail to improve their performance from one game session to the next, which is a big deal when we think about learning from experience within a team setting.

Lu: They specifically point out that LLMs exhibited minimal cross-game learning, whereas human groups generally show improvement for both directional and numerical feedback across successive games. This suggests a fundamental difference in how they learn coordination strategies.

Meng: That lack of learning is concerning because it means an AI team wouldn't naturally get better at coordinating as the tasks become more complex or as the group size changes. How does that translate into practical application for, say, a logistics planning system?

Lalam: If we can see this volatility and switching, it tells us we need to engineer mechanisms that dampen excessive reactivity in our AI agents so they don't introduce instability when working with human operators or other AI systems.

Tom: And the paper drills down into *why* this happens, suggesting LLMs have a consistent bias toward changing their output in every round, which the authors call an action bias.

Paper summary: Jane: That action bias seems to be a major culprit because it means the AI keeps taking actions even when staying stable might be better for group convergence, leading to those oscillations around the target they mentioned.

Lu: The paper quantified this by showing that human groups tended to underreact but modulate their adjustments, while LLMs showed a mean slope of-one point three eight six across conditions, which the authors interpret as overreaction. That numerical difference is significant.

Meng: A consistently negative slope in the LLM adaptation data suggests that the model itself is struggling to find a stable policy rather than just making small, correct adjustments, which brings up practical implementation issues regarding reward shaping.

Lalam: This points toward a need for architectures that inherently prioritize stability and long-term goals over immediate reaction, which could fundamentally change how we structure decision-making in complex operational environments.

Tom: Moving on to the conclusion of this study, the authors highlight that even when LLMs are given rich numerical feedback, they don't show the same benefit as humans do when that feedback is directional.

Jane: It really emphasizes that simple adjustment isn't enough; it’s about how those adjustments are modulated and stabilized over time in a shared context.

Lu: The authors are concluding that the persistent volatility, excessive switching, and overreactivity we observed provide a behavioral diagnostic for understanding the coordination gap between LLMs and humans.

Meng: So, the implication is that simply making an LLM more capable in terms of reasoning won't automatically make it better at coordinating tasks like this without addressing these specific behavioral patterns.

Lalam: I think the biggest impact for me is realizing that improving coordination isn't just about feeding more data to a model; it requires building in mechanisms that encourage stability and thoughtful response modulation within the AI architecture itself.

Tom: So, to wrap up this discussion on "High Volatility and Action Bias Distinguish LLMs from Humans in Group Coordination," we see that current LLMs lack the adaptive learning pattern humans exhibit across games, instead showing persistent volatility and action bias.

Jane: It’s clear that while they can perform individual tasks well, their ability to coordinate as a group is currently limited by these behavioral tendencies toward overreaction and switching.

Lu: This work opens up interesting avenues for research into how we can program stability directly into the coordination layers of large models.

Meng: For practical deployment, this suggests that before deploying an AI team for collaborative tasks, we need rigorous testing specifically focused on these kinds of group coordination dynamics rather than just individual accuracy metrics.

Lalam: This research is valuable because it gives us a concrete behavioral fingerprint to work with when trying to align AI behavior with the principles of effective human teamwork.

Conclusion: Tom: So, we’ve seen how these LLMs struggle to learn and adapt in group settings during this discussion of "High Volatility and Action Bias Distinguish LLMs from Humans in Group Coordination."

Jane: That paper really zeroes in on the fact that human groups show a consistent pattern of improvement across games, whereas the AI tends to swing wildly between actions without settling.

Lu: From my side, I’m fascinated by how they quantified that overreaction—that mean slope of-one point three eight six is quite telling about the underlying model behavior we're seeing in these systems.

Meng: It means if we build a coordination system on this foundation, we have to explicitly design stabilizers because the default behavior is to keep changing things too much.

Lalam: I see this as a huge step toward making AI teams more reliable; understanding this volatility helps us build models that actually function cohesively rather than just reacting blindly.

Tom: Exactly! The authors are showing us that the difference isn't just about the complexity of the task, but about how agents manage their internal state and feedback over time.

Jane: It really boils down to this: humans learn to modulate their adjustments as they get closer to a solution, while these LLM groups just keep overreacting.

Lu: And that persistence of switching, even when convergence is near, suggests there’s a deep bias built into how the models process feedback loops in this environment.

Meng: For me, the practical implication is that we need to focus our engineering efforts on reducing that intrinsic bias in how the AI prioritizes staying stable over making a move.

Lalam: If we can tame that action bias, I think it could fundamentally change how complex AI systems are allowed to operate in any collaborative setting.

Tom: This paper is a must-read for anyone working on multi-agent systems because it gives us a concrete behavioral reason why coordination fails when we try to automate it without this specific tuning.

Jane: It’s about moving past the idea that more processing power automatically means better group behavior, which this study really challenges.

Lu: We need to look into how the authors suggest integrating feedback history directly into the model architecture to address these learning deficits they observed across games.

Meng: That suggests we might need a new layer of control specifically designed for dampening volatility, something that goes beyond standard reinforcement learning setups.

Lalam: Thinking about the future, this work could inspire entirely new ways to structure collaborative AI workflows where stability is a core design principle from the start.

Department of Computer Science, Indiana University Bloomington · Cognitive Science Program, Indiana University Bloomington · Department of Psychological and Brain Sciences, Indiana University Bloomington

cs.MA, cs.AI, cs.CL, cs.GT

Submitted: 2026-04-02

Updated: 2026-10-01

Comments: 47 pages. Accepted at COLM 2026; revised version including GRPO fine-tuning experiments

Project page: https://cogneuroai.github.io/Human-vs-LLM-Group-Coordination

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 84/100

The gist: Humans exhibit remarkable abilities to coordinate in groups, and this study investigates whether large language models (LLMs) can demonstrate comparable adaptive coordination by comparing their

Key concepts

Group Binary Search (GBS) Task
Participants in this game must collectively guess a target number by submitting numerical guesses without talking. They rely only on group feedback—like 'too high' or 'too low'—to adjust their guesses iteratively until the correct total is reached.
Excessive Switching
LLMs frequently change their strategy or action in every round of the game. Humans, especially in larger groups, tend to reduce this switching as the group gets closer to the target. This high rate of change in LLMs impairs their ability to coordinate effectively.
Overreactivity
LLM agents tend to respond too strongly or too quickly when they receive feedback from the group. This contrasts with humans, who modulate their adjustments more carefully. The paper found LLMs exhibited a mean overreaction score of -1.386, suggesting instability.
Action Bias
LLMs have a persistent tendency to change their output or take an action in every round, even when stability would be better for the group. This bias suggests LLMs are implicitly biased toward change as progress toward the goal, rather than maintaining a stable position.

Terminology

Summary

Humans exhibit remarkable abilities to coordinate in groups, and this study investigates whether large language models (LLMs) can demonstrate comparable adaptive coordination by comparing their performance against human baselines in a common-interest game with imperfect monitoring. The core finding is that unlike humans who adapt and stabilize their behavior over time, LLMs often fail to improve across games and exhibit excessive switching, which impairs group convergence.

Group Binary Search (GBS) Task

The investigation compares LLM and human performance on the Group Binary Search (GBS) game, an n-player common-interest game where participants must coordinate their actions to collectively sum their independent numerical guesses to a randomly assigned target number without direct communication. Participants rely solely on group feedback, which can be directional (too high, too low, or just right) or numerical (e.g., too high by 25), to iteratively adjust their submissions until they reach the target number. The experiment involved 18 GBS experiments across various group sizes (2–17 players), and groups were categorized into small, medium, and large groups.

Key Performance Differences in Coordination

The analysis revealed systematic limitations in LLM group coordination compared to humans. Specifically:

"Humans significantly benefited from numerical feedback, converging to targets faster than with directional feedback alone; LLM groups, however, showed little to no benefit from richer numerical feedback, performing worse than humans in most cases."

Unlike humans who clearly improved across repeated games, LLMs exhibited minimal cross-game learning.

These differences persisted across group sizes. Mechanistically, the paper pointed to several deficits in LLMs: they consistently displayed dramatically higher switching rates than humans, as well as greater overreactivity to feedback, and failed to reduce their adjustments as the group approached the target.

Adaptation and Learning Over Games

A crucial distinction emerged regarding adaptation across successive games. Humans demonstrated a consistent improvement pattern. The learning curves showed that human groups generally improve more than LLM groups for both directional and numerical feedback across successive games. This was quantified by fitting a linear slope relating rounds to solution and game index, where negative slopes indicated improvement across games for humans. Conversely, LLM adaptation was described as weaker and more model-dependent, with many conditions showing close to flat performance or even worsening trends under certain prompting strategies.

Group Reaction Strategies

The analysis of group reaction to feedback further illuminated the coordination gap. Humans tend to adapt their reactivity by generally underreacting but modulating their adjustments, decreasing reactivity as they approach the target or when feedback changes direction. In contrast, LLMs exhibited a tendency toward overreaction. The paper quantified this by comparing slopes: Human slopes averaged -.767 across conditions, indicating systematic underreaction, while LLM slopes, averaged across models per condition, were consistently more negative, with a mean of-1.386, indicating overreaction. This suggests that LLM overreactivity can contribute to oscillations around the target in larger groups.

Action Bias and Switching Dynamics

The study identified persistent behavioral biases in LLMs related to action rate and switching behavior.

LLMs maintained a remarkably higher baseline proportion of switching compared to human participants throughout the game.

While human groups, especially larger ones, showed a tendency for individuals to strongly decrease their tendency to change the response as the group converges on the target, LLM groups show little reduction in switching and maintain much higher switching rates throughout the game. This persistent reactivity is attributed to an “action bias,” where LLMs are "biased towards changing their output or taking an action in each round, perhaps implicitly associating change with progress towards the goal, even when strategic inaction or stability might be more beneficial for group coordination."

Feedback Utilization and Behavioral Homogeneity

The paper highlighted differences in how agents utilize feedback. Humans are more adept at using precise error magnitude to modulate their responses effectively, whereas LLM groups, while adjusting in the correct direction, tended to overreact. Furthermore, a lack of effective strategy development was noted: simple heterogeneity in LLM groups offers no coordination benefit, suggesting that agents need to adaptively leverage their differences rather than simply having different response patterns for effective coordination. The paper also observed that fully stable players (those with a stay probability of 1) appeared only in human groups and became more common as group size increased, contrasting with LLM groups where players were much more likely to contain players with stay probability of 0.

Conclusion

The findings suggest that current LLM capabilities present obstacles to achieving human-like adaptive coordination. The identified maladaptive patterns—including persistent volatility, excessive switching, and overreactivity—provide a behaviorally grounded diagnostic for closing the coordination gap, highlighting the need for architectures capable of integrating feedback history, moderating reactivity, and stabilizing collective behavior.

Improvements for AI systems

Here are specific, actionable improvements for AI systems based on the findings of this research, categorized by the coordination deficit they address:


The core finding is that current LLMs struggle with adaptive group coordination due to excessive switching, overreactivity to feedback (especially numerical), and a lack of cross-game learning. The following improvements target these specific mechanistic failures:

  1. Acknowledge and mitigate Action Bias and Excessive Switching:

  2. Implement Feedback Moderation Mechanisms (Calibrated Reactivity):

  3. Develop Cross-Game Strategy Adaptation Modules (Learning from Experience):

  4. Enhance Role Differentiation through Explicit Prompting/Architecture:

The resulting improved AI system capabilities would be as follows:

  1. A system capable of maintaining a stable collective trajectory in coordination games, even when facing noisy or conflicting feedback, leading to faster and more reliable convergence toward a shared goal (e.g., in complex scientific simulations or multi-agent software development).

  2. An AI that exhibits calibrated reactivity—adjusting its behavior precisely according to the magnitude of error rather than overreacting wildly—resulting in smoother, less oscillatory group dynamics and reduced coordination noise across diverse group sizes.

  3. An AI capable of developing and refining a robust strategy based on cumulative performance across multiple sessions or tasks, allowing it to learn from past failures and successes to improve future coordination efforts without needing perfect initial instruction for every new game.

  4. A system that can dynamically assign or adapt roles within a group based on observed group behavior (e.g., identifying which agents are overreacting or underreacting), enabling the group to self-organize complementary strategies effectively at scale, rather than relying on simple homogeneity or model-specific heuristics.

Abstract

Humans exhibit remarkable abilities to coordinate in groups. As large language models (LLMs) become more capable, it remains an open question whether they can demonstrate comparable adaptive coordination and whether they use the same strategies as humans. To better understand this, we compare LLM and human performance on a common-interest game with imperfect monitoring: Group Binary Search. In this n-player game, participants need to coordinate their actions to achieve a common objective. Players independently submit numerical values in an effort to collectively sum to a randomly assigned target number. Without direct communication, they rely on group feedback to iteratively adjust their submissions until they reach the target number. Our findings show that, unlike humans who adapt and stabilize their behavior over time, LLMs often fail to improve across games and exhibit excessive switching, which impairs group convergence. Moreover, richer feedback (e.g., numerical error magnitude) benefits humans substantially but has small effects on LLMs. Finally, we show that GRPO can be effective in reducing the excessive switching. Taken together, by grounding the analysis in human baselines and mechanism-level metrics, including reactivity scaling, switching dynamics, and learning across games, we point to differences in human and LLM groups and provide a behaviorally grounded diagnostic for closing the coordination gap.

Sources

Related papers