Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games

arXiv:2505.24255 · cs.CL, cs.AI, cs.HC · Submitted 2026-08-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games".

Jane: The paper was written by Neemesh Yadav, Palakorn Achananuparp, Jing Jiang and Ee-Peng Lim from Singapore Management University and Australian National University and (Note: The affiliations are linked to the authors' names, but the organizations themselves are listed here.).

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary of Methodology: Tom: The core of this research involves running two thousand seven hundred simulations across six different LLM models to see how they perform when facing these specific social pressures in the Ultimatum Game.

Jane: They tested a wide variety of reasoning methods, ranging from simple Chain-of-Thought or CoT, against those more complex methods involving Theory of Mind (ToM).

Meng: The setup is incredibly detailed; they simulate multiple rounds of negotiation until a specific stopping condition is met or the maximum turns are reached.

Lu: It's interesting that the researchers found that having both ToM levels—zero-order introspection combined with first-order reasoning—often yields more impactful results than just relying on one level alone.

Lalam: And it’s clear from the data that when agents were initialized with "Fair" beliefs, they tended to behave in a way humans expect them to, showing great consistency across the entire two thousand seven hundred-game set.

Tom: The methodology is solid, but I think we need to appreciate how complex they made this by ensuring each agent had its own internal reasoning process.

Jane: They also used different roles—proposer and responder—which adds a layer of strategic complexity that’s crucial for realism in the dynamic negotiation.

Lu: The way they structured the BDI model allows us to see exactly how an agent moves from what it believes another agent knows, to what it intends to do next.

Meng: From an engineering standpoint, this confirms that building a a dedicated reasoning loop around the LLM is far more effective than just making it part of the prompt structure.

Lalam: It shows us exactly how we can structure our future AI agents—not just as black boxes, but as entities with internal logic and social awareness.

Improvements and Contributions: Tom: The study is really making a lot of contributions by introducing this comprehensive framework for studying agentic behavior against human norms in these simulations.

Jane: It’s not just one level of ToM they are testing; the authors are expertly testing combinations like Selfless-Fair, which allows us to see how specific beliefs interact with strategic decisions across different agents.

Meng: I think the biggest contribution lies in quantifying these behavioral improvements using metrics like Deviation Scores (DS) and measuring acceptance rates rather than just looking at raw outputs.

Lu: The fact that they are comparing those specific prosocial beliefs—Greedy, Fair, Selfless—is a huge theoretical contribution to the idea that an agent' drive personality drives its strategic outcomes in the game.

Lalam: This entire framework improves our ability to design AI agents that aren’t just statistically correct in their language, but genuinely capable of understanding and reacting appropriately within complex social dynamics.

Tom: That combination of measuring deviation from human expectation is what really sets this apart, it moves beyond just a simple pass or fail.

Jane: And it shows that the specific interplay between two roles—like Greedy Proposer facing a Selfless Responder—can lead to very different negotiation outcomes.

Lu: The theoretical contribution is that the belief state becomes an explicit driver of strategy, which is a huge step toward making those real-world applications possible.

Meng: It allows us to build tools and systems where we can predict exactly how an agent will behave based on its internal programming, which is critical for reliability in financial or social contexts.

Lalam: This work creates a blueprint for designing AI that has genuine social awareness, ensuring it doesn' that its actions are predictable and align with human norms.

Conclusion: Tom: So, we’ve seen how Theory of Mind reasoning and these prosocial beliefs significantly impact the alignment of LLMs in negotiation tasks, proving that internal strategy matters.

Jane: It’s clear from the results that when you want to see the most rational behavior, having those agents adopt a Fair-Fair belief combination is what really works best for matching human expectations.

Lu: I think this work suggests that the future of AI isn't just about building bigger models, but fundamentally about crafting smarter internal reasoning architectures.

Meng: We can now start building practical systems that incorporate these specific cognitive modules to improve how AI behaves in complex tasks like negotiation or resource allocation in the real world.

Lalam: To conclude on this entire paper "Effects of Theory of Mind and Prosocial Belief on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games," it is a powerful demonstration that the path toward trustworthy, human-aligned AI involves deeply understanding its internal reasoning capacity.

Tom: It’s truly an exciting time to watch as we see how AI can reason about intentions.

Jane: It’s a huge relief to see the results showing that align with human expectations.

Lu: I'm just thrilled by the possibility of seeing this concept applied across different scenarios.

Meng: I'm looking forward to building these practical modules into our systems, too.

Lalam: It feels like we are finally moving towards an era where AI can genuinely understand and engage with human intentions.

Conclusion: Tom: So we've seen how Theory of Mind reasoning and prosocial beliefs significantly impact the alignment of LLMs in negotiation tasks, proving that internal strategy matters for a fair outcome.

Jane: It’s clear from the results that when you want to see the most rational behavior, having those agents adopt a Fair-Fair belief combination is what really works well against human expectations.

Lu: I think this work suggests that the future isn't just about building bigger models, but fundamentally about creating smarter internal reasoning architectures that guide our AI.

Meng: We can now start building practical systems that incorporate these specific cognitive modules to improve how AI handles complex tasks like negotiation or resource allocation in the real world.

Lalam: I view this as a powerful step, seeing the ability to create trustworthy, human-aligned AI by understanding its internal reasoning capacity within the a controlled environment of Theory of Mind and Prosocial Belief.

Tom: That is exactly what I mean; it moves us away from simply having AI generate text and toward genuine social competence.

Jane: It’s encouraging to see the results so that we can be more certain when building these agents, knowing they are more likely to behave in a predictable, human-like way.

Lu: This really opens up possibilities for thinking about how we might integrate these models into complex, multi-agent simulations later on.

Meng: We're already seeing the path to making those operationalized systems that actually function reliably in the real world, which is what matters most from an engineering standpoint.

Lalam: Let's wrap up this discussion with a final thought on the paper "Effects of Theory of Mind and Prosocial Belief on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games," because its findings truly illuminate how far we have come.

Tom: I hope that will be our last topic for today, as we've got a fascinating new paper coming up that’s exploring the nuances of trust in AI.

Neemesh Yadav, Palakorn Achananuparp, Jing Jiang, Ee-Peng Lim

Singapore Management University · Australian National University · (Note: The affiliations are linked to the authors' names, but the organizations themselves are listed here.)

cs.CL, cs.AI, cs.HC

Submitted: 2026-08-24

Updated: 2026-08-25

Code: https://github.com/Stealth-py/UltimatumToM

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 72/100

The gist: The study investigates "the role of ToM reasoning on the performances of different LLMs in negotiation tasks" to determine if incorporating Theory of Mind (ToM) results in agent behaviors that are

Key concepts

Theory of Mind (ToM)
A complex reasoning method tested in the simulations. ToM allows agents to model what another agent knows or intends to do next, moving beyond simple responses. This capability adds a crucial layer of strategic complexity necessary for realistic negotiation tasks.
Prosocial Beliefs
These are specific internal moral beliefs (such as 'Fair,' 'Selfless,' or 'Greedy') assigned to AI agents. The research uses these beliefs to test how an agent's programmed sense of duty or morality drives its strategic decisions and outcomes within the game.
LLM Alignment
This refers to designing Large Language Models so that their behavior is predictable and matches human expectations. The study suggests that achieving alignment requires moving beyond simple text generation by incorporating dedicated internal reasoning modules.

Terminology

Summary

The study investigates the role of ToM reasoning on the performances of different LLMs in negotiation tasks to determine if incorporating Theory of Mind (ToM) results in agent behaviors that are more closely aligned with human norms.

Research Gap and Objectives:

A significant research gap is addressed regarding how strategic reasonings, such as ToM, can be used to steer behavioral alignment of LLMs in social simulations. The study aims to examine the role of ToM reasoning on the performances of different LLMs in negotiation tasks using the ultimatum game (UG) as a controlled environment.

Methodology and Setup:

The researchers utilized a multi-level ToM reasoning framework combined with various prosocial beliefs. Agents were initialized with three specific prosocial beliefs: Greedy, Fair, and Selfless. These beliefs determine the range of proposed splits and decisions in the game.

The agents were tested across six diverse LLMs: GPT-4o-mini, GPT-4o, o3-mini [a reasoning model], DeepSeek R1 Distilled Qwen 32B [a reasoning model], Llama 3.3-70B, and Llama 3.1-8B.

The study employed four distinct reasoning methods:

  1. Chain-of-thought (CoT)

  2. Zero-order ToM (Introspection)

  3. First-order ToM (Inferring the other agent’s mental state)

  4. Both (Combining zero+first order)

The experiment involved 9 possible combinations of Proposer and Responder beliefs, resulting in a total of 2,700 simulated games across all models. The game setup involves a single-stake, multiple-round format (up to 10 rounds), allowing agents to engage in negotiations spanning multiple rounds.

** Evaluation Metrics:**

The agents were evaluated using two sets of metrics:

  1. Performance Metrics: Acceptance Rate (AC), Average number of turns/rounds required to finish a game (Avg. Turns), and Total Payout for the proposer (lambda P) and responder (lambda R).

  2. Behavioral Metrics: Deviation Scores (DS) which measure deviation from expected human behavior:

  • P: Deviation of initially proposed shares (Proposer perspective).

  • RA: Deviation of finally accepted shares (Responder perspective).

  • RR: Deviation of rejected shares from expected accepted shares (Responder perspective).

Key Findings:

The results, derived from the 2,700 simulations, indicate that ToM reasoning enhances behavior alignment, decision-making consistency, and negotiation outcomes.

Specific findings include:

  • Belief Types: The Fair-Fair belief combination achieves the highest overall AC, closely followed by the Fair-Selfless combination. Economically, the rational behavior of a responder is to accept any non-zero share; however, due to beliefs, this rationality is challenging to determine.

  • Model and Reasoning Impact:

  • For Proposers, GPT-4o (β=-0.493) and o3-mini (β=-0.442) produced significantly lower deviation scores, indicating that GPT-4o achieves the most significant least deviations when acting as a Proposer.

  • For Responders, Llama 3.1, being the smallest model, achieves the most significant (β=-0.4537) least deviation from our expectations in Responder Accepted Shares.

  • Optimal Strategy: The study found that first-order ToM was effective for proposers, while combined ToM improved responder’s acceptance decisions. Furthermore, rejecting offers was best handled by simpler reasoning: reject[ing offers] implies that simpler (zero-order or CoT) reasoning is better than complex methods (first-order or both ToM).

  • Prosocial Alignment: The most aligned scenarios were found when the Proposer and Responder were both Fair.

Conclusion:

The study concludes that ToM reasoning significantly enhances behavior alignment compared to other reasoning methods, providing a nuanced understanding of how diverse belief types and ToM reasoning contribute to negotiation dynamics in human-AI interaction.

Improvements for AI systems

Based on the rigorous findings presented in this study, I have identified several critical areas where current AI agents can be significantly improved to achieve human-aligned strategic behavior in complex social simulations and negotiations.

We must move beyond simple Chain-of-Thought (CoT) prompting and integrate a multi-layered reasoning architecture that models the opponent's mental state, as demonstrated by the success of ToM (Theory of Mind) in this research.

1. Integrated Theory of Mind (ToM) Module:

  • Improvement: Implement a dedicated, mandatory module within the agentic pipeline that forces explicit inference about the other player’s Belief (B), Desires (D), and Intentions (I). This module must operate before the final decision or proposal is generated.

  • Mechanism: Forcing agents to answer questions like: What does Player B believe? (as seen in the 'First-order ToM' prompts) allows the agent to calculate a probability of acceptance based on that inferred state.

  • Capability: The improved system can now simulate strategic adaptation, moving beyond simple maximizing utility toward predictive social alignment.

2. Belief/Personality Conditioning Layer:

  • Improvement: Agents must be initialized not just with a static persona, but with a quantifiable set of Prosocial Belief parameters (Greedy, Fair, Selfless). This acts as a weight or bias on the decision-making function.

  • Mechanism: The system uses the specific belief set (e.g., 'Fair') to constrain the possible outcomes (e.g., only proposing offers near 50%).

  • Capability: The improved system can exhibit consistent and predictable behavioral patterns, making its actions reliably attributable to a defined psychological profile, rather than random output.

3. Dynamic Reasoning Selection (Contextual Switching):

  • Improvement: Implement a dynamic logic gate that selects the appropriate reasoning level based on the agent's role and the current interaction phase.

  • Mechanism:

  • For Proposers: Utilize First-order ToM, allowing them to predict how a specific belief (e.g., 'Fair') will react to an offer, leading to higher alignment with human norms (beta = -0.493 for GPT-4o).

  • For Responders: Use Both ToM when accepting a complex proposal (to maximize acceptance) and utilize Zero-order ToM/CoT when rejecting an offer (to maintain efficiency and minimize complexity, as suggested by the strong performance in rejection scenarios).

  • Capability: The improved system can execute contextually optimal reasoning, adapting its cognitive load to the specific requirements of a social interaction.


By implementing these architectural changes, the resulting AI system will demonstrate superior performance in strategic simulations:

1. Enhanced Behavioral Alignment (The Fair Standard):

  • The system will naturally gravitate toward Fair-Fair outcomes as its default or most robust strategy, achieving the highest Acceptance Rate (AC) and lowest Deviation Score (DS). This ensures that the AI aligns with the empirical human preference for equitable outcomes in negotiation.

2. Strategic Role Optimization:

  • The system can execute Proactive Negotiation: As a Proposer, it will use ToM to craft an opening offer that is optimally designed to be accepted by a target belief (e.g, targeting a 'Fair' responder with an equal split), rather than simply maximizing its own share.

  • The system can execute Adaptive Response: As a Responder, it will not just reject based on fixed rules, but will calculate the probability of rejection based on the inferred beliefs of the proposer, leading to more sophisticated counter-offers or strategic acceptance.

3. Increased Efficiency and Consistency:

  • By utilizing Zero-order ToM for rejections, the system can achieve high efficiency (low Average Turns) while maintaining a consistent behavioral profile, reducing unnecessary conversational loops that plague unguided LLM agents.

Summary of Improved System Capability: The improved AI system will be able to navigate dynamic negotiations with social intelligence, making decisions that are not just mathematically optimal, but socially plausible, mirroring the complex interplay between belief and strategy observed in human behavior.

Abstract

Large Language Models (LLMs) have shown potential in simulating human behaviors and performing theory-of-mind (ToM) reasoning, crucial for complex social interactions. We investigate ToM reasoning's role in aligning agentic behaviors with human norms in negotiation tasks, using the ultimatum game as our referenced task. We initialized LLM agents with different prosocial beliefs (Greedy, Fair, Selfless) and reasoning methods (chain of thought and ToM reasoning of varying levels), examining their decision-making process and outcome across multiple LLMs, including reasoning models like o3-mini and DeepSeek-R1 Distilled Qwen 32B. We perform 2,700 simulations to show that ToM reasoning enhances behavioral alignment with human, decision-making consistency, and negotiation outcomes. Consistent with prior findings, reasoning LLMs exhibit limited capability compared to ToM-enhanced LLMs, with different game roles benefiting from different ToM orders. Fair proposers and responders accepting offers were the most consistent with their strategic reasonings, whereas all agents showed strong consistencies with human beliefs when rejecting offers, except when the offer was fair. Human verification further revealed that Llama 3.3 70B produces reasoning most consistent with its actions and beliefs. Our findings advance understanding of ToM's role in human-AI interaction and cooperative decision-making. The code used for our experiments can be found at https://github.com/Stealth-py/UltimatumToM.

Sources

Related papers