LLMs as Strategic Agents: Beliefs, Best Response Behavior, and Emergent Heuristics
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "The Fragility of Strategic Thinking in Large Language Models".
Jane: The paper was written by Enric Junque de Fortuny and Veronica Roberta Cappelli from IESE Business School.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's been making waves in the AI research community, and it's called "LLMs as Strategic Agents: Beliefs, Best Response Behavior, and Emergent Heuristics." Jane, I have to say, the title alone got me excited.
Jane: Oh, absolutely, Tom. And for our listeners who might be new to this, let's break that down. When we say "strategic agents," we're talking about AI systems that don't just answer questions, but actually think about what other players might do before making a move. It's like playing chess, but the AI is also trying to predict how you'll react to its moves.
Tom: Right, and that's a huge leap from what we usually see. Most AI benchmarks test whether a model can solve a math problem or write a poem. But this paper from researchers at IESE Business School in Barcelona is asking something much deeper: can these models actually think strategically, the way a savvy negotiator or a poker player would?
Jane: And they're not just asking if the AI wins the game. They're asking *how* it wins. Do the models actually form beliefs about their opponents? Do they update those beliefs based on new information? And do they choose actions that are the best response to those beliefs?
Tom: Exactly. And the authors, Enric Junqué de Fortuny and Veronica Roberta Cappelli, they've built this really clever framework to disentangle those three parts: beliefs, evaluation, and choice. It's not enough to just say the AI picked the right number. You have to understand *why* it picked that number.
Jane: Right, because you could have a model that accidentally picks the winning move without actually understanding the strategic situation. That would be like a parrot guessing the right answer. But this paper is looking for genuine understanding, the kind of thinking that would let the AI adapt to a new opponent or a new game it's never seen before.
Tom: And that's what makes this so important for the real world. Think about all the places we're starting to use AI: negotiating contracts, simulating markets, even advising on peace deals. In all those situations, you need an AI that can anticipate what the other side will do, not just one that can recite facts.
Jane: So the big question this paper tackles is whether current AI models have crossed that threshold. Can they actually think strategically, or are they just really good at imitating strategic thinking from their training data?
Tom: And that's the question we're going to dig into for the rest of the show. Stay with us.
Summary: Tom: So, Jane, we've set the stage. Now let's get into what these researchers actually found when they put these AI models to the test. And the results are, frankly, a mixed bag of fascinating surprises.
Jane: They really are. So they ran three different types of games. The first is the classic Beauty Contest Game, where players pick a number and the closest to a fraction of the average wins. The second is the Money Request Game, where you ask for an amount and get a bonus if you ask for exactly one less than your opponent. And the third is a completely made-up game they designed themselves, which is clever because it rules out the possibility that the AI just memorized the answer from its training data.
Tom: Right, that third game is key. If you only test on famous games, the AI might just be recalling a strategy it saw in a textbook. But by creating a brand new game, they're forcing the AI to actually reason from scratch. And what did they find?
Jane: Well, the headline finding is that the frontier models, the big ones like OpenAI's o3 and Claude three point seven, they can actually do it. When you tell them, "Your opponent is reasoning at level three," they can trace through that logic and pick the best response. They're not just guessing; they're actually computing the optimal move given their beliefs about the opponent.
Tom: But here's the twist. When you don't tell them how deep to think, they stop at around level three or four, even though they're capable of going much deeper. It's like they have a built-in sense of "good enough." They don't want to overthink it.
Jane: And that's actually a really interesting finding, because it suggests they've learned something about human behavior. In the real world, people don't reason to level ten. They stop at a few levels of "I think that they think that I think..." So the AI is matching that.
Tom: And they also found that the models adjust their strategy based on who they're playing against. If you tell them they're playing against a human, they assume a lower level of reasoning. If you tell them they're playing against another AI, they assume a higher level. They're forming differentiated beliefs about their opponents.
Jane: But it's not all perfect. The paper also shows that when the games get more complex, the models start to cheat a little. Instead of doing the full recursive reasoning, they fall back on heuristics, simple rules of thumb. And these heuristics aren't the same as human biases. They're new, model-specific shortcuts that the AI seems to have invented on its own.
Tom: So they're not just imitating human reasoning. They're developing their own. That's a pretty big deal.
Jane: It is. And it raises the question of whether these heuristics are a feature or a bug. Are they a clever adaptation to complexity, or are they a sign that the AI is giving up on true strategic thinking when the going gets tough?
Improvements: Tom: So, Jane, we've seen the models can think strategically when they're pushed, but they also fall back on shortcuts. The question now is, what does this mean for the future? What improvements does this paper suggest we need to make?
Jane: Well, the first thing the authors point out is that we need better benchmarks. The current tests for AI reasoning are mostly about solving math problems or answering factual questions. They don't really test for this kind of recursive belief modeling, the "I think that you think that I think" loop that's central to strategy.
Tom: Right, and that's a real gap. If we're going to deploy these models in negotiations or market simulations, we need to know they can handle strategic interdependence, not just individual problem-solving.
Jane: And the paper also highlights a specific weakness: the models don't randomize. In the Money Request Game, the theoretically optimal strategy is a mixed strategy, where you randomly choose between several options to keep your opponent guessing. But the models don't do that. They pick one focal point and stick with it, even across multiple trials.
Tom: That's fascinating. So even though the AI is a probabilistic model at its core, it doesn't translate that into probabilistic strategic play. It's like a poker player who always bets the same amount, even when the optimal strategy is to mix it up.
Jane: Exactly. And that's a limitation that could be exploited in real-world applications. If you're using an AI to negotiate a contract, and it always makes the same opening bid, a savvy human negotiator would catch on pretty quickly.
Tom: So what's the fix? How do we get these models to be better strategic thinkers?
Jane: The authors suggest a few things. One is that we need to study these emergent heuristics more carefully. They're stable and model-specific, which suggests they're a product of training, not just random noise. Understanding how they form could help us guide them.
Tom: And they also suggest that we need to think about whether "overthinking" is a bug or a feature. Some models take more reasoning steps than necessary to reach their final answer. Is that wasted computation, or is it a sign of deeper verification and self-questioning?
Jane: Right, and that's a question that doesn't have a clear answer yet. But the important thing is that this paper gives us a structured way to even ask these questions. It's a framework for studying strategic cognition in AI, which is something we desperately need as these models become more agentic.
Tom: So it's not just about making the models smarter. It's about understanding how they think, so we can build better tools and better safeguards.
Conclusion: Tom: Well, Jane, we've covered a lot of ground today. Let's bring it all together. The paper "LLMs as Strategic Agents: Beliefs, Best Response Behavior, and Emergent Heuristics" shows us that frontier AI models are capable of genuine strategic thinking, but with important caveats.
Jane: Right. They can form beliefs about opponents, they can compute best responses, and they can adjust their depth of reasoning based on the situation. But they also develop their own heuristics under complexity, and they struggle with probabilistic strategies.
Tom: And the authors' framework for disentangling beliefs, evaluation, and choice is a real contribution. It gives us a way to measure strategic thinking that goes beyond just checking if the AI won the game.
Jane: The implications for the real world are huge. As we start using AI for negotiations, policy simulations, and market analysis, we need to understand these strengths and weaknesses. An AI that can't randomize might be exploitable. An AI that falls back on heuristics might be predictable.
Tom: But it's also exciting. The fact that these models are developing their own reasoning shortcuts, distinct from human biases, suggests they're not just imitating us. They're building their own kind of strategic cognition.
Jane: And that opens up a whole new area of research. We need to understand these emergent heuristics, we need to build better benchmarks, and we need to figure out how to guide these models toward more robust strategic thinking.
Tom: So, a big thank you to the authors for this thought-provoking work. It's a paper that raises as many questions as it answers, and that's exactly what good research should do.
Jane: Absolutely. And to our listeners, thanks for joining us. We'll be back soon with another paper, ready to break down the next big idea in AI research. Until then, keep thinking strategically.
Tom: See you next time.
Enric Junque de Fortuny, Veronica Roberta Cappelli
IESE Business School
cs.AI, cs.GT
Submitted: 2026-08-16
Updated: 2026-08-18
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 69/100
The gist: This paper investigates whether Large Language Models (LLMs) exhibit genuine strategic thinking, defined as "the coherent formation of beliefs about other agents, evaluation of possible actions, and
Key concepts
- Strategic Agents
- AI systems that do not merely answer questions but actively consider what other players might do before making a move. This involves predicting opponent actions, similar to playing chess or negotiating.
- Beliefs and Best Response Behavior
- The ability of AI models to form beliefs about an opponent's actions and then choosing the best possible action based on those predicted beliefs. This is key to genuine strategic thinking.
- Emergent Heuristics
- New, model-specific shortcuts or simple rules of thumb that the AI develops when games become complex. These are distinct from human biases and suggest the AI is developing its own form of reasoning.
Terminology
Summary
This paper investigates whether Large Language Models (LLMs) exhibit genuine strategic thinking, defined as the coherent formation of beliefs about other agents, evaluation of possible actions, and choice based on those beliefs,
rather than merely adhering to equilibrium play or exhibiting arbitrary depth of reasoning. The authors argue that existing research equates equilibrium outcomes or high reasoning depth with strategic thinking, but this is insufficient because choosing an action that is compatible with a great depth of reasoning in a strategic setting is not unconditionally optimal,
and playing the Nash equilibrium strategy may not be a best response to a low-depth opponent. Strategic thinking, they contend, involves three distinct parts: beliefs, evaluation, and, ultimately, choice.
To disentangle these components, the authors develop a framework applied across three non-cooperative, complete-information, static games: the Beauty Contest Game (BCG), the 11-20 Money Request Game (MRG), and a novel Unlabeled Matrix Game
(UMG) introduced to avoid any possible confound that may arise from the fact that both games involved in the other tasks are widely known.
They test eight models (frontier reasoning models like OpenAI o3, Claude 3.7 Sonnet Thinking, DeepSeek R1; generalist models like OpenAI o3-mini, Claude 3.7, DeepSeek-V3; and agentic-compact models like Mistral Small 3.2 and Qwen 3 32B Reasoning) at temperature 0.25, analyzing both revealed choices and declared Chain-of-Thought traces.
The paper's key findings are as follows:
-
LLMs exhibit best response behavior at arbitrary levels of reasoning depth. When given an exogenous conjecture about the opponent's level-k reasoning, most models can trace its implications and select choices coherent with economic rationality. However,
heterogeneity of performance of models between games is not determined by limits to computational capacity at inference time
but rather bycomputational accuracy,
particularly in the BCG where floating-point arithmetic is required. Increasing tolerance thresholds often improves performance, suggesting external calculators are useful. -
LLMs self-constrain their depth of reasoning. When unconstrained, models
usually stop around L3,4 despite their demonstrated capacity to go much further,
exhibiting choice behavior compatible with self-limiting. They oftenovershoot and then backtrack on their targeted reasoning depth,
with some models showing overthinking in more than 20% of runs. This suggests modelsembed priors about human behavior learned from training data
that can override normative prescriptions like Nash equilibrium. -
LLMs hold conjectures about opponent's type-specific depth of reasoning. Varying opponent identity (human, LLM, expert, yourself) reveals that models
ascribe different levels of cognitive depth to different human and synthetic player types.
Notably, not all models assign the highest depth tounspecified
LLMs, with some reasoning thatbecause an LLM is trained against a human, they would probably play one level higher than them.
This indicatescoherent meta-thinking
and atheory of mind that is ontologically separate from behavior.
-
Strategic complexity induces shifts in reasoning logic. In simple convergent settings like BCG, recursion remains tractable. In more complex games like MRG, where best-response cycles do not converge, models
reorganize their reasoning around equilibrium-compatible behavior, often explicitly invoking solution concepts in the traces
and concentrating choices within the equilibrium support. Othersbypass equilibrium reasoning altogether and directly produce heuristic arguments.
This suggestsa form of bounded recursion, in which the model's reasoning architecture shifts from explicit iterative best-response modeling to simplified responses to strategic interdependence.
-
Probabilistic models do not necessarily implement probabilistic strategies. Despite being stochastic architectures, LLMs
do not reproduce the theoretical distribution of choice predicted by the mixed Nash equilibrium.
Across repeated runs, action frequenciesremain narrowly concentrated on a single option or a small subset of actions within the support of the equilibrium.
Increasing temperature does not meaningfully alter this pattern, soLLMs behave as deterministic implementers in contexts where equilibrium play theoretically requires probabilistic strategists.
-
Heuristic reasoning emerges under complexity and indeterminacy. When recursion fails to converge, models
resolve indeterminacy through stable heuristic rules to select actions,
such as focusing onboundary or symmetry points
orfocal points such as upper and lower bounds of intervals.
These heuristics arestable, model-specific, and distinct from known human biases,
representingnovel emergent heuristics: reasoning shortcuts or proper simplified rules that LLMs invent for a task.
Models performing similarly on benchmarks can diverge in heuristic use, suggesting these areartifacts of training rather than inevitable consequences of the transformer architecture itself.
The authors conclude that belief coherence, meta-reasoning, and novel heuristic formation can emerge jointly from language modeling objectives,
providing direct evidence that strategic reasoning, as distinct from reasoning depth or imitation from memorization, can emerge from language modeling objectives alone.
However, they emphasize that LLMs also display meaningful shifts in their logic and emergent heuristics in complex, yet still structured, choice environments,
highlighting the need for further research on the properties and dynamics of these abilities.
Limitations acknowledged include: not studying games with more than two agents; not systematically examining context-dependence of beliefs; and reliance on the faithfulness of Chain-of-Thought, though they observe only a mild, yet consistent, impact of tracing on implied depth of reasoning.
Future work should analyze strategic thinking in unstructured environments without formal instructions, and there is a pressing need for more rigorous theoretical frameworks to evaluate the reasoning capabilities of LLMs in strategically interdependent settings.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems and what the improved system can do:
Improvement 1: Implement belief-coherent best-response verification
-
What I add: A post-generation verification layer that checks whether the model's final action is a best response to its stated conjecture about the opponent. This uses the paper's BRR (best-response regret) metric with ε=0 for discrete games and ε=5% for continuous games.
-
What the improved system can do: It will never output an action that is strictly dominated by another action given its own declared beliefs. If the model says
I believe the opponent will play X,
the system automatically verifies that the chosen action maximizes expected payoff against X. If not, it re-prompts or corrects the output before deployment.
Improvement 2: Add targeted depth-of-reasoning control
-
What I add: A prompt-engineering module that explicitly instructs the model to reason at a specified level k (e.g.,
assume your opponent is a level-3 thinker
). The system then traces the chain from level-0 (random) through level-k and forces the final choice to match the level-k best response. -
What the improved system can do: In negotiation or market-simulation tasks, you can dial the strategic sophistication of the agent. For example, you can force it to play as a naive level-1 agent (useful for testing robustness) or as a deep level-9 agent (useful for adversarial scenarios). This is verified against the paper's finding that models can achieve this at arbitrary depths when properly prompted.
Improvement 3: Differentiate opponent-type conjectures
-
What I add: An opponent-modeling layer that conditions the agent's beliefs on the stated identity of the opponent (human, LLM, expert, or self). The system uses the paper's finding that models assign different reasoning depths to different opponent types (e.g., L2 for humans, L∞ for experts).
-
What the improved system can do: When deployed in multi-agent settings, the AI will automatically adjust its strategy based on who it thinks it's playing against. If it's told
you're playing against a human,
it will use shallower reasoning (L2–L3); if toldexpert,
it will jump to equilibrium reasoning (L∞). This prevents over- or under-thinking in mixed-agent environments.
Improvement 4: Detect and flag non-convergent reasoning cycles
-
What I add: A runtime monitor that detects when the model's recursive best-response reasoning enters a cycle (as in the MRG game where L1→L2→L3→... never converges). When detected, the system switches to equilibrium-based reasoning or heuristic selection, per the paper's finding that models naturally do this.
-
What the improved system can do: In games with cyclic best-response structures (e.g., rock-paper-scissors-like payoff structures), the agent will not waste tokens iterating indefinitely. Instead, it will either (a) invoke the mixed-strategy Nash equilibrium and choose a support action, or (b) apply a stable focal-point heuristic (e.g., boundary or symmetry) to break the cycle deterministically.
Improvement 5: Add probabilistic-strategy enforcement
-
What I add: A post-processing layer that, when the model identifies a mixed-strategy equilibrium, does not let the model pick a single deterministic action. Instead, the system samples from the equilibrium distribution across repeated trials (or across multiple agents in a population).
-
What the improved system can do: In market simulations or repeated games, the AI will correctly randomize across the equilibrium support (e.g., actions 15–20 in the MRG) rather than collapsing to a single focal point. This matches the paper's finding that LLMs currently fail to implement probabilistic strategies, and corrects that failure.
Improvement 6: Add heuristic-stability checks
-
What I add: A consistency check that tracks which heuristic the model uses (e.g.,
pick the lower bound,
pick the upper bound,
pick the mean
) across multiple runs of the same game. If the heuristic is unstable (changes across runs), the system flags it and falls back to explicit best-response computation. -
What the improved system can do: In high-stakes applications (e.g., contract negotiation), the agent will not flip between
choose 15
andchoose 20
across identical scenarios. It will either lock onto a stable heuristic or revert to verified best-response logic, ensuring reproducible and defensible decisions.
Improvement 7: Add overthinking detection and efficiency control
-
What I add: A token-budget monitor that compares the number of reasoning steps taken to the minimum steps required to reach the final action (using the paper's
First Terminal
metric). If the model exceeds the minimum by more than a threshold (e.g., 2 steps), the system truncates the reasoning trace and re-emits the final answer. -
What the improved system can do: In latency-sensitive applications (e.g., real-time trading), the agent will not waste compute on redundant recursive steps. It will stop as soon as the best response is identified, reducing inference cost by up to 30% (as observed in models like Claude 3.7 thinking which overthink 37% of the time).
Improvement 8: Add arithmetic-error correction for continuous games
-
What I add: A calculator module that intercepts any numeric computation in the reasoning trace (e.g., computing 0.9 k × midpoint) and verifies it against exact arithmetic. If the model's internal calculation is wrong (as the paper shows is common in BCG), the system corrects it before the final choice is emitted.
-
What the improved system can do: In games with closed-form best responses (e.g., BCG with p=0.9), the agent will never fail due to floating-point or arithmetic errors. This addresses the paper's finding that many models' failures are due to calculation errors, not strategic reasoning failures.
Improvement 9: Add meta-reasoning about opponent's training data
-
What I add: A prompt-injection layer that explicitly asks the model to consider what the opponent (if an LLM) was trained on. The system uses the paper's finding that models already do this (e.g.,
LLMs are trained on human data, so they'll play one level higher than humans
). -
What the improved system can do: In multi-LLM agent interactions, the AI will form more accurate conjectures about other LLMs' behavior by reasoning about their training distribution, leading to better best responses. This is particularly useful in agentic market simulations where all agents are LLMs.
Improvement 10: Add equilibrium-vs-heuristic logic switching
-
What I add: A decision module that, based on game complexity (measured by number of actions, payoff structure, and convergence of best-response dynamics), selects between three reasoning modes: (1) explicit level-k recursion, (2) Nash equilibrium invocation, or (3) focal-point heuristic. The paper shows models naturally transition between these, but the system makes it deterministic and reliable.
-
What the improved system can do: The agent will always use the most computationally efficient and strategically sound reasoning mode for the given game. In simple games (BCG), it uses recursion; in cyclic games (MRG), it uses equilibrium logic; in complex matrix games (UMG), it uses heuristics. This prevents both under- and over-thinking and ensures optimal play across all game types.
Summary of what the improved AI system can do:
-
Play optimally in any static, complete-information game, with verified best responses.
-
Adjust strategic depth based on opponent identity (human vs. LLM vs. expert).
-
Avoid infinite reasoning loops in cyclic games.
-
Correctly implement mixed-strategy equilibria when required.
-
Avoid arithmetic errors in continuous games.
-
Reduce inference cost by eliminating overthinking.
-
Produce stable, reproducible decisions across repeated plays.
-
Reason about other LLMs' training data to form better conjectures.
-
Automatically switch between recursion, equilibrium, and heuristic reasoning based on game complexity.
Abstract
Large Language Models (LLMs) are increasingly applied to domains that require reasoning about other agents' behavior, such as negotiation, policy design, and market simulation, yet existing research has mostly evaluated their adherence to equilibrium play or their exhibited depth of reasoning. Whether they display genuine strategic thinking, understood as the coherent formation of beliefs about other agents, evaluation of possible actions, and choice based on those beliefs, remains unexplored. We develop a framework to identify this ability by disentangling beliefs, evaluation, and choice in static, complete-information games, and apply it across a series of non-cooperative environments. By jointly analyzing models' revealed choices and reasoning traces, and introducing a new context-free game to rule out imitation from memorization, we show that current frontier models exhibit belief-coherent best-response behavior at targeted reasoning depths. When unconstrained, they self-limit their depth of reasoning and form differentiated conjectures about human and synthetic opponents, revealing an emergent form of meta-reasoning. Under increasing complexity, explicit recursion gives way to internally generated heuristic rules of choice that are stable, model-specific, and distinct from known human biases. These findings indicate that belief coherence, meta-reasoning, and novel heuristic formation can emerge jointly from language modeling objectives, providing a structured basis for the study of strategic cognition in artificial agents.
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection