How a Cooperative-Override Circuit Suppresses Nash Play in Large Language Models
cs.GT, cs.AI, cs.LG
Submitted: 2026-04-29
Updated: 2026-09-18
Comments: v3: major revision. Title changed (previously "What Suppresses Nash Equilibrium Play in Large Language Models? Mechanistic Evidence and Causal Control"). Main text rewritten at 12 pages; mechanistic campaign re-run under a seeded, hash-verified protocol; new 48-game payoff-random experiment; several earlier-version claims corrected, with all protocol changes documented in Appendix H
License: http://creativecommons.org/licenses/by/4.0/
The gist: On the named Prisoner's Dilemma under direct prompting, three larger instruction-tuned models, Llama-3-70B, Qwen2.5-32B, and Qwen2.5-72B, lock at full cooperation, the metric's maximum distance from
Terminology
Abstract
On the named Prisoner's Dilemma under direct prompting, three larger instruction-tuned models, Llama-3-70B, Qwen2.5-32B, and Qwen2.5-72B, lock at full cooperation, the metric's maximum distance from Nash with zero variance across replicates, while Llama-3-8B plays near-Nash. Opening the models, a logit-lens analysis finds a distributed cooperative override. Intermediate readouts lean toward the Nash action through roughly three quarters of network depth before a late surge toward cooperation, and the final layer settles the contest. The size of that final correction, not the surge, rank-matches chain-of-thought behavior across scale and two architectures. In the 8B the override is a single causally controllable direction in the residual stream; steering it dials the decision, and clamping its component at one position of one layer moves the choice strictly monotonically, Spearman rho = 1.000, with generation fluent. The circuit is lexical. It survives name removal and payoff rescaling but disengages when Cooperate and Defect are replaced with neutral labels, and on 48 payoff-random games with neutral surfaces no model locks cooperative on any dilemma or shows general equilibrium competence. In mixed-model populations a single Nash-playing agent collapses cooperation contagiously. What suppresses Nash play in large language models is a word-triggered circuit rather than missing competence, and it can be measured, bounded, and controlled.
Sources
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Persona Vectors in Games: Measuring and Steering Strategies via Activation Vectors
- Steering Language Models With Activation Engineering
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Exact Regret Frontiers and Externality Scheduling in Centralized Serial-Dictatorship Bandits
- In-Context Credit Assignment via the Core
- Breaking 1/epsilon Barrier in Quantum Zero-Sum Games: Generalizing Metric Subregularity for Spectraplexes
- Enhancing Affine Maximizer Auctions with Correlation-Aware Payment
- LLM Bidders Preserve the Mechanism-Level Orderings of Human Bidders
- Towards Performatively Stable Equilibria in Decision-Dependent Games for Arbitrary Data Distribution Maps