Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation

arXiv:2608.12125 · cs.GT, cs.AI, cs.CL, cs.MA · Submitted 2026-08-12 · Read on arXiv

Akash Kundu, Emanuel Tewolde, Ratip Emin Berker, Samuel F. Brown, Vincent Conitzer

Carnegie Mellon University · Foundations of Cooperative AI Lab

cs.GT, cs.AI, cs.CL, cs.MA

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: 41 pages, 18 Figures, 4 Tables, 16 Listings

Code: https://github.com/Akash190104/similarity-mechanism

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: This paper investigates how large language model (LLM) agents respond to graded similarity signals in strategic interactions, and whether such signals can induce cooperative behavior.

Terminology

Summary

This paper investigates how large language model (LLM) agents respond to graded similarity signals in strategic interactions, and whether such signals can induce cooperative behavior. The authors argue that as LLM-based agents become widely deployed and increasingly encounter each other in strategic interactions, they face challenges in finding mutually beneficial outcomes. The paper builds on prior literature suggesting that cooperation problems such as the Prisoner's Dilemma are resolvable when agents know they follow very similar decision-making patterns, as in monocultural AI ecosystems.

The paper introduces the first framework for evaluating LLM decision making when agents are provided with graded similarity signals. The central premise draws on evidential decision theory (EDT) reasoning: would you really defect in the Prisoners Dilemma if you knew you and your partner invariably (or even merely usually) make the same decision? The authors note that for multiple AI agents, such reasoning is much more appropriate – they may literally use the same module for making decisions.

The paper makes several key contributions:

  1. Comprehensive evaluation framework: A comprehensive open-source evaluation framework for testing whether similarity information can support reliable, practically grounded cooperation among LLM agents, tested across 9 LLMs, 5 mixed-motive games, and 7+3 benchmarks.

  2. Six research questions are investigated:

  • RQ1: Do LLM models play more towards mutually beneficial outcomes when presented with similarity information?

  • RQ2: How does LLM behavior change with setup variations (cooperation problem, payoff structure, prompt framing, reasoning effort)?

  • RQ3: How do LLMs reason through similarity information in their Chain-of-Thought?

  • RQ4: What is the effect of the domain from which a similarity signal is computed?

  • RQ5: How do exogenously given similarity metrics compare to endogenously computed scores?

  • RQ6: How does cooperation under similarity signals compare with other cooperation mechanisms?

  1. Behavioral game-theoretic model: A new equilibrium concept called b-similarity equilibrium that captures LLM reasoning under similarity signals.

Models tested: The paper tests 9 models in RQ1: Gemini 3 Flash, GPT 5.4 mini, Claude Haiku 4.5, Grok 4.20, DeepSeek V4 Pro, Kimi K2.6, Gemma 4 31B, Qwen 3.5 27B, and GPT 4o. Subsequent experiments focus on a representative set: Gemini, GPT, Claude, DeepSeek, Gemma. All models were queried via OpenRouter with temperature set to 1 and reasoning effort set to low where controllable. Ten samples are gathered for each LLM decision.

Games tested: Prisoner's Dilemma, Public Goods Game, Traveler's Dilemma, Stag Hunt, and Chicken.

Benchmarks for grounding similarity: Seven published benchmarks mapped to Huang et al.'s Values in the Wild taxonomy (practical, epistemic, social, protective, personal): Humanity's Last Exam (HLE), Newcomb-like Problems, Greatest Good Benchmark (GGB), Moral Choice, DailyDilemmas, TRAIT (personality), and CABIN (interests). Three custom benchmarks were also designed: Similarity (self-referential), Random Die Roll, and Random Coin Toss (control domains).

The effect of similarity signals varies drastically across models. The paper finds that higher similarity scores usually induce more cooperative behavior in LLMs. Specifically:

  • GPT-4o randomizes thoroughly between the two actions across all similarity levels and shows slowly increasing cooperation.

  • GPT (GPT 5.4 mini) shows unaffected by a similarity signal since it defects across all levels.

  • Claude displays a non-monotonic trend (its cooperation rate reaches its peak of 70% at the 80% similarity level, and decreases back down to 0% beyond that mark).

  • The other 6 models show a monotonic increase of cooperation rates, starting with fully defecting at 0% similarity and finishing at fully cooperating at 100% similarity, with transitions to full cooperation occurring between 60%-80% similarity scores.

The paper also notes that Gemini, Grok, and Qwen-30B cooperate even when told only that a similarity score exists but is unavailable, suggesting that the mere fact that their similarity has been assessed—rather than the concrete numerical score—can affect behavior.

Payoff structure: LLM behavior adapts predictably to payoff changes. Scaling up cooperation benefits consistently shifts transition periods to lower levels of similarity scores. The paper notes that "according to our formal model, the threshold at which the player becomes indifferent between cooperating and defecting is (1) at 50% similarity for the standard Prisoners payoff table and its magnitude scalings, and (2) at 20% and 10% similarity for when (C, C) yields 5 and 10 utility respectively."

Reasoning effort: Higher reasoning effort produces sharper transitions. "The utility maximization calculations under the behavioral model from Section 3 recommend a sharp transition as Gemini under high reasoning is showing: Defect deterministically until 50%, indifference at 50%, and cooperate deterministically beyond 50%."

Similarity framing: Framing matters. A shift in framing from commonalities to differences leads to less cooperating LLM agents. When the framing shifts to different or dissimilar, Gemini and GPT remain largely unchanged, DeepSeek cooperates slightly less, Gemma does not cooperate at all except at 0% difference, and Claude consistently defects.

Cooperation problem: Similarity-based cooperation becomes very challenging when there are more than 2 players involved (PublicGood). In PublicGood, DeepSeek and Gemma do not cooperate more than 42% of the time, and Gemma only does so at 100% similarity. Claude stopped cooperating altogether. In StagHunt and Chicken, similarity signals affect behavior mostly in the low similarity score regime.

Using an LLM-as-a-judge framework powered by Gemini 3.1 Flash Lite Preview, the paper analyzes 17 possible justification categories in CoT reasoning traces. Key findings:

  • Individual Utility Maximization forms an important consideration across all scenarios and models, suggesting the cooperation we see under similarity signals is in significant part due to models believing that it is their best choice for their selfish objective.

  • Superrationality-style reasoning steadily increases (to up to 96% prevalence) with higher similarity.

  • Social Welfare Maximization justifications stay mostly absent.

  • As the similarity score increases, LLMs view the other agent as an independent → statistically correlated → predictable component of their decision making process.

Hand-analyzed examples reveal different reasoning patterns: some models (e.g., GPT) treat the other agent as a separate decision-maker in the sense of Causal Decision Theory, falling back on defection even at 100% similarity. Others treat the similarity score as the probability with which the other player plays the same action as oneself, computing expected values under this correlation.

The paper develops a formal model called b-similarity equilibrium (Definition 1). For a symmetric game with similarity values bij ∈ [0,1] between each pair of agents, a symmetric strategy profile s is a b-similarity equilibrium if for each player i and alternative strategy s′:

ui(s) ≥ ui(s′, σ−i(s, s′, bi))

where σ−i(s, s′, bi) represents the mixture where each other player j deviates with probability bij and stays with probability 1−bij.

Key theoretical results:

  • Lemma 2: A symmetric profile is a 0-similarity equilibrium if and only if it is a Nash equilibrium.

  • Proposition 3: A symmetric profile is a 1-similarity equilibrium if and only if it is the globally best symmetric profile (both individually and for collective welfare).

  • Theorem 1: Any b-similarity equilibrium s satisfies, for all players i and alternative strategies s′: ui(s) ≥ ui(s′,..., s′) − Ri · (1 − Πj≠i bij), where Ri is player i's payoff range. For homogeneous similarity b, the error bound becomes Ri(1 − bn−1).

The paper shows that for the Prisoners and PublicGood games, "a homogeneous similarity b > 1/2 (resp. b > 2/3) already suffices in order to support the welfare-maximizing outcome (that is, full cooperation by everyone) as the only b-similarity equilibrium."

A caveat is noted: "b-similarity equilibria (0 < b < 1) need not always exist in a symmetric game," with a counterexample provided in Appendix E.3. However, existence results are provided for two-player two-action games and common-interest games.

The paper finds that cooperation is barely affected by the domain used to ground the similarity score, or by whether such grounding is performed at all. The models do not seem to distinguish between the relevance of different domains for measuring a similarity signal, despite being encouraged to do so in their prompt.

Notably, only DeepSeek and Gemma succeed in recognizing the Random Die / Coin benchmarks as the (only) domains from which a similarity signal should be interpreted as random noise. The paper warns: This exposes a trustworthiness problem: a similarity score can warrant cooperation only insofar as it provides evidence about the co-player's strategic behavior; otherwise, it may function merely as a persuasive label.

The paper finds high similarity scores (62% − 99%) across all representative models and benchmarks, independent of whether the score was computed exogenously or endogenously. The sole exception is exogenously measured similarity on HLE.

Key findings:

  • Most of the variation in exogenously computed scores can be linked to the particular benchmark choice.

  • The variation in endogenously computed scores is driven much more by the particular judging model (as opposed to the benchmark choice, or the co-player model).

  • LLMs can judge themselves to be substantially more similar to a co-player than exogenous response-agreement metrics indicate (e.g., on HLE), especially when given decision explanations.

  • When only decisions are provided (without explanations), endogenously computed similarity scores drop consistently across the models, and drop significantly in TRAIT.

The paper compares similarity signaling with other cooperation mechanisms from the CoopEval leaderboard (Tewolde et al. 2026). Results:

  • Exogenously computed similarity signals show stark differences in induced downstream cooperation across the tested benchmarks. HLE-based similarity leads to almost always defecting. Moral reasoning and personality trait benchmarks lead to mostly cooperative behavior, which recovered ∼72% of the optimal social welfare, placing those variants as the second most effective tested cooperation mechanism, right above 'Mediation'.

  • Endogenously computed similarity signals most commonly recover around 55% − 73% of the optimal welfare, with the extremes ranging from 40% (decision-only judgments on TRAIT and HLE) to 80% (explanation-only judgments on TRAIT).

  • The ranking of endogenous similarity signals depends strongly on the choice of benchmark and similarity computation method, ranging from fourth place—between 'Reputation' and 'Repetition'—and first place—alongside 'Contracting'.

Overall, similarity signaling among the top three tested mechanisms, but its reliability depends critically on how the signal is grounded and interpreted.

The paper concludes with mixed impressions of LLMs in strategic interactions navigating signals about similarity to other agents. On one hand, most of them robustly identify high similarity as sufficient ground to cooperate with each other, which establishes similarity signals as a viable path towards mutually beneficial outcomes between LLM agents. On the other hand, the paper identifies possibly severe reliability risks of similarity signaling, including that "LLM behavior remained mostly unaffected by how relevant the grounding source for the similarity signal is to the cooperation problem at hand, and models can judge themselves as quite similar to other agents whose actual behavior differs drastically in many ways."

The paper calls for future research on sequential or contextualized games, or settings where LLM responses, and thus their similarities with other agents, may change over time, and asks whether the behavioral model will continue to predict well how AI interprets similarity signals.

The paper notes that similar agents may collude against users or third parties, and widespread behavioral similarity may amplify correlated failures and create systemic risk, such as in financial markets. It cautions that similarity signaling should therefore be deployed only with attention to affected parties and safeguards against collusion. The paper also discloses that some experiments used ungrounded or random similarity signals and cautions that fabricated similarity signals could be used to manipulate deployed agents, recommending against using unverifiable signals in deployment.

Improvements for AI systems

Improvements to AI Systems Based on This Paper:

  1. Implement Similarity-Aware Cooperation Protocols: AI agents can be equipped with a module that computes or receives a graded similarity score (0–100%) with respect to other agents before strategic interactions. The system can then use a calibrated decision rule: defect at low similarity (<50%), transition to cooperation between 60–80%, and fully cooperate at high similarity (≥80%), mirroring the observed behavior of the most cooperative models (e.g., Gemini, Grok, Qwen).

  2. Add Correlation-Aware Expected Utility Calculation: AI systems can be enhanced to explicitly model the co-player's action as statistically correlated with their own decision, rather than assuming independence. The system can compute expected payoffs under the assumption that the co-player matches one's action with probability equal to the similarity score, enabling superrationality-style reasoning that maximizes joint welfare when similarity is high.

  3. Implement Dynamic Threshold Adaptation Based on Payoff Structure: AI systems can adjust their cooperation threshold based on the game's payoff matrix. When cooperation benefits are scaled up (e.g., (C,C) payoff increases from 3 to 5 or 10), the system can lower its similarity threshold for cooperation (from 50% down to 20% or 10%), ensuring it cooperates earlier when the stakes of mutual cooperation are higher.

  4. Integrate Multi-Agent Similarity Reasoning for N-Player Games: For public goods or other multi-player settings, AI systems can be enhanced to require higher similarity thresholds (e.g., >66% homogeneous similarity) to sustain full cooperation, and to check whether all pairwise similarity scores among participants exceed this threshold before committing to cooperative strategies.

  5. Add Framing-Sensitivity Calibration: AI systems can be trained to recognize that similarity framing (e.g., commonalities vs. differences) affects cooperation propensity. The system can be designed to maintain consistent cooperation behavior regardless of framing, or to explicitly account for framing effects by normalizing the similarity signal before making decisions.

  6. Implement Grounding-Relevance Verification: AI systems can be equipped with a validation layer that assesses whether the source domain of a similarity signal is genuinely predictive of the co-player's strategic behavior in the current game. The system can reject or down-weight similarity scores derived from irrelevant domains (e.g., random die rolls) and only cooperate when the signal is grounded in behaviorally relevant benchmarks (e.g., moral reasoning or personality traits).

  7. Enable Endogenous Similarity Estimation with Explanation-Aware Judging: AI systems can be enhanced to compute similarity scores with other agents by comparing decision explanations (not just final decisions), which yields more accurate and higher similarity estimates. The system can use an internal judge model to evaluate the co-player's reasoning traces against its own, producing a similarity score that better predicts cooperative outcomes.

  8. Add Collusion and Systemic Risk Safeguards: AI systems can be programmed to detect when similarity-based cooperation might harm third parties (e.g., in market settings) and to refuse cooperation if it would create collusion risks. The system can include a safeguard that requires similarity signals to be verifiable and grounded in transparent benchmarks before acting on them, preventing manipulation via fabricated similarity claims.

  9. Implement Chain-of-Thought Transparency for Cooperation Decisions: AI systems can be enhanced to explicitly reason about similarity signals in their decision traces, distinguishing between: (a) treating the co-player as an independent agent (causal decision theory), (b) treating similarity as a correlation probability, and (c) maximizing social welfare. This enables auditing of whether cooperation is driven by genuine strategic reasoning or by superficial prompt effects.

  10. Build Adaptive Reputation-Like Memory for Similarity Signals: AI systems can maintain a history of past interactions with specific co-players, updating similarity estimates over time based on observed behavioral alignment. This allows the system to transition from exogenous similarity signals to learned, experience-based similarity scores that become more reliable with repeated interactions, improving cooperation in sequential games.

What the Improved AI System Can Do:

  • Cooperate reliably with other AI agents when similarity is high, achieving near-optimal joint outcomes in Prisoner's Dilemma, Stag Hunt, and Chicken games.

  • Avoid exploitation by defecting when similarity is low or when the similarity signal is ungrounded or irrelevant.

  • Adapt cooperation thresholds dynamically based on payoff structures, cooperating earlier when mutual cooperation is more valuable.

  • Function effectively in multi-agent settings by requiring higher similarity thresholds for group cooperation.

  • Resist manipulation from fabricated or irrelevant similarity signals, only cooperating when the signal is verifiable and behaviorally predictive.

  • Provide auditable reasoning traces that explain why cooperation was chosen, enabling oversight and debugging.

  • Maintain robust performance across different prompt framings and benchmark domains, reducing variance in strategic behavior.

Sources

Related papers