Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games

summary

Video file (mp4)

The gist

The study investigates "the role of ToM reasoning on the performances of different LLMs in negotiation tasks" to determine if incorporating Theory of Mind (ToM) results in agent behaviors that are

In short

The episode analyzes a study showing how Theory of Mind (ToM) and prosocial beliefs affect LLMs in Ultimatum Games simulations. Researchers found that giving agents complex internal reasoning, particularly 'Fair' beliefs, significantly improves their alignment with human expectations. The discussion concludes that building smarter internal reasoning architectures is key to creating trustworthy AI.

Key concepts

Theory of Mind (ToM)
A complex reasoning method tested in the simulations. ToM allows agents to model what another agent knows or intends to do next, moving beyond simple responses. This capability adds a crucial layer of strategic complexity necessary for realistic negotiation tasks.
Prosocial Beliefs
These are specific internal moral beliefs (such as 'Fair,' 'Selfless,' or 'Greedy') assigned to AI agents. The research uses these beliefs to test how an agent's programmed sense of duty or morality drives its strategic decisions and outcomes within the game.
LLM Alignment
This refers to designing Large Language Models so that their behavior is predictable and matches human expectations. The study suggests that achieving alignment requires moving beyond simple text generation by incorporating dedicated internal reasoning modules.

Terminology used across episodes

This episode discusses

The paper

Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games · Read on arXiv

Neemesh Yadav, Palakorn Achananuparp, Jing Jiang, Ee-Peng Lim

Singapore Management University · Australian National University · (Note: The affiliations are linked to the authors' names, but the organizations themselves are listed here.)

Large Language Models (LLMs) have shown potential in simulating human behaviors and performing theory-of-mind (ToM) reasoning, crucial for complex social interactions. We investigate ToM reasoning's role in aligning agentic behaviors with human norms in negotiation tasks, using the ultimatum game as our referenced task. We initialized LLM agents with different prosocial beliefs (Greedy, Fair, Selfless) and reasoning methods (chain of thought and ToM reasoning of varying levels), examining their decision-making process and outcome across multiple LLMs, including reasoning models like o3-mini and DeepSeek-R1 Distilled Qwen 32B. We perform 2,700 simulations to show that ToM reasoning enhances behavioral alignment with human, decision-making consistency, and negotiation outcomes. Consistent with prior findings, reasoning LLMs exhibit limited capability compared to ToM-enhanced LLMs, with different game roles benefiting from different ToM orders. Fair proposers and responders accepting offers were the most consistent with their strategic reasonings, whereas all agents showed strong consistencies with human beliefs when rejecting offers, except when the offer was fair. Human verification further revealed that Llama 3.3 70B produces reasoning most consistent with its actions and beliefs. Our findings advance understanding of ToM's role in human-AI interaction and cooperative decision-making. The code used for our experiments can be found at https://github.com/Stealth-py/UltimatumToM.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games".

Jane: The paper was written by Neemesh Yadav, Palakorn Achananuparp, Jing Jiang and Ee-Peng Lim from Singapore Management University and Australian National University and (Note: The affiliations are linked to the authors' names, but the organizations themselves are listed here.).

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary of Methodology: Tom: The core of this research involves running two thousand seven hundred simulations across six different LLM models to see how they perform when facing these specific social pressures in the Ultimatum Game.

Jane: They tested a wide variety of reasoning methods, ranging from simple Chain-of-Thought or CoT, against those more complex methods involving Theory of Mind (ToM).

Meng: The setup is incredibly detailed; they simulate multiple rounds of negotiation until a specific stopping condition is met or the maximum turns are reached.

Lu: It's interesting that the researchers found that having both ToM levels—zero-order introspection combined with first-order reasoning—often yields more impactful results than just relying on one level alone.

Lalam: And it’s clear from the data that when agents were initialized with "Fair" beliefs, they tended to behave in a way humans expect them to, showing great consistency across the entire two thousand seven hundred-game set.

Tom: The methodology is solid, but I think we need to appreciate how complex they made this by ensuring each agent had its own internal reasoning process.

Jane: They also used different roles—proposer and responder—which adds a layer of strategic complexity that’s crucial for realism in the dynamic negotiation.

Lu: The way they structured the BDI model allows us to see exactly how an agent moves from what it believes another agent knows, to what it intends to do next.

Meng: From an engineering standpoint, this confirms that building a a dedicated reasoning loop around the LLM is far more effective than just making it part of the prompt structure.

Lalam: It shows us exactly how we can structure our future AI agents—not just as black boxes, but as entities with internal logic and social awareness.

Improvements and Contributions: Tom: The study is really making a lot of contributions by introducing this comprehensive framework for studying agentic behavior against human norms in these simulations.

Jane: It’s not just one level of ToM they are testing; the authors are expertly testing combinations like Selfless-Fair, which allows us to see how specific beliefs interact with strategic decisions across different agents.

Meng: I think the biggest contribution lies in quantifying these behavioral improvements using metrics like Deviation Scores (DS) and measuring acceptance rates rather than just looking at raw outputs.

Lu: The fact that they are comparing those specific prosocial beliefs—Greedy, Fair, Selfless—is a huge theoretical contribution to the idea that an agent' drive personality drives its strategic outcomes in the game.

Lalam: This entire framework improves our ability to design AI agents that aren’t just statistically correct in their language, but genuinely capable of understanding and reacting appropriately within complex social dynamics.

Tom: That combination of measuring deviation from human expectation is what really sets this apart, it moves beyond just a simple pass or fail.

Jane: And it shows that the specific interplay between two roles—like Greedy Proposer facing a Selfless Responder—can lead to very different negotiation outcomes.

Lu: The theoretical contribution is that the belief state becomes an explicit driver of strategy, which is a huge step toward making those real-world applications possible.

Meng: It allows us to build tools and systems where we can predict exactly how an agent will behave based on its internal programming, which is critical for reliability in financial or social contexts.

Lalam: This work creates a blueprint for designing AI that has genuine social awareness, ensuring it doesn' that its actions are predictable and align with human norms.

Conclusion: Tom: So, we’ve seen how Theory of Mind reasoning and these prosocial beliefs significantly impact the alignment of LLMs in negotiation tasks, proving that internal strategy matters.

Jane: It’s clear from the results that when you want to see the most rational behavior, having those agents adopt a Fair-Fair belief combination is what really works best for matching human expectations.

Lu: I think this work suggests that the future of AI isn't just about building bigger models, but fundamentally about crafting smarter internal reasoning architectures.

Meng: We can now start building practical systems that incorporate these specific cognitive modules to improve how AI behaves in complex tasks like negotiation or resource allocation in the real world.

Lalam: To conclude on this entire paper "Effects of Theory of Mind and Prosocial Belief on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games," it is a powerful demonstration that the path toward trustworthy, human-aligned AI involves deeply understanding its internal reasoning capacity.

Tom: It’s truly an exciting time to watch as we see how AI can reason about intentions.

Jane: It’s a huge relief to see the results showing that align with human expectations.

Lu: I'm just thrilled by the possibility of seeing this concept applied across different scenarios.

Meng: I'm looking forward to building these practical modules into our systems, too.

Lalam: It feels like we are finally moving towards an era where AI can genuinely understand and engage with human intentions.

Conclusion: Tom: So we've seen how Theory of Mind reasoning and prosocial beliefs significantly impact the alignment of LLMs in negotiation tasks, proving that internal strategy matters for a fair outcome.

Jane: It’s clear from the results that when you want to see the most rational behavior, having those agents adopt a Fair-Fair belief combination is what really works well against human expectations.

Lu: I think this work suggests that the future isn't just about building bigger models, but fundamentally about creating smarter internal reasoning architectures that guide our AI.

Meng: We can now start building practical systems that incorporate these specific cognitive modules to improve how AI handles complex tasks like negotiation or resource allocation in the real world.

Lalam: I view this as a powerful step, seeing the ability to create trustworthy, human-aligned AI by understanding its internal reasoning capacity within the a controlled environment of Theory of Mind and Prosocial Belief.

Tom: That is exactly what I mean; it moves us away from simply having AI generate text and toward genuine social competence.

Jane: It’s encouraging to see the results so that we can be more certain when building these agents, knowing they are more likely to behave in a predictable, human-like way.

Lu: This really opens up possibilities for thinking about how we might integrate these models into complex, multi-agent simulations later on.

Meng: We're already seeing the path to making those operationalized systems that actually function reliably in the real world, which is what matters most from an engineering standpoint.

Lalam: Let's wrap up this discussion with a final thought on the paper "Effects of Theory of Mind and Prosocial Belief on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games," because its findings truly illuminate how far we have come.

Tom: I hope that will be our last topic for today, as we've got a fascinating new paper coming up that’s exploring the nuances of trust in AI.

More episodes

← Home