Assessing mentalization in humans and large language models
cs.AI, q-bio.NC
Submitted: 2026-08-26
Updated: 2026-08-26
Comments: 46 pages, 4 figures, 2 tables. Supplementary information available on request from the authors
License: http://creativecommons.org/licenses/by/4.0/
The gist: Mentalization - the ability to infer others' beliefs and intentions to guide one's own choices - is a key cognitive function underlying human social interactions.
Terminology
Abstract
Mentalization - the ability to infer others' beliefs and intentions to guide one's own choices - is a key cognitive function underlying human social interactions. Large language models (LLMs) demonstrate behaviour consistent with humans on theory-of-mind tasks, yet whether these models can guide adaptive behaviour through mentalization is unknown. Here we use two economic games with cognitive computational modeling to uncover the latent strategies underlying mentalization in LLMs. We tested individual LLM agents across four model families, DeepSeek, GPT-4.1, GPT-5 and Gemini 2.0 Flash (N = 2,099), against opponents of varying sophistication and examined whether a prompting strategy designed to elicit strategic reasoning improved performance. We benchmarked results against human participants (N = 251) as a comparative measure. Across both games, LLMs showed clear behavioural and computational signatures of mentalizing that differed markedly by model provider and size. Strategic prompting generally improved performance by inducing more sophisticated reasoning, yet the extent of the benefit differed across the two tasks. Last, GPT-5 agents flexibly adapted their recursive depth of reasoning to increasingly sophisticated opponents, demonstrating superior performance to human participants. Collectively, we demonstrate different capacities for mentalization across LLMs, and highlight cognitive computational modeling as a formal method for assessing comparative intelligence across humans and machines.
Sources
- AI Behavioral Science
- Machine Psychology
- Moral Foundations of Large Language Models
- MoCa: Measuring Human-Language Model Alignment on Causal and Moral Judgment Tasks
- Strategic Chain-of-Thought: Guiding Accurate Reasoning in LLMs through Strategy Elicitation
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks
- How FaR Are Large Language Models From Agents with Theory-of-Mind?
- Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
- Do Large Language Models Have a Planning Theory of Mind? Evidence from MindGames: a Multi-Step Persuasion Task
- Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models
- Re-evaluating Theory of Mind evaluation in large language models
- Levels of Analysis for Large Language Models
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- Position: Theory of Mind Benchmarks are Broken for Large Language Models
- Mind Your Theory: Theory of Mind Goes Deeper Than Reasoning
- Do Theory of Mind Benchmarks Need Explicit Human-like Reasoning in Language Models?
- A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios
- LLM Strategic Reasoning: Agentic Study through Behavioral Game Theory
- Strategic Behavior of Large Language Models: Game Structure vs. Contextual Framing
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection