MARBO: Relational Belief Grounding for LLM Agents in Social Deduction Games
cs.AI, cs.LG
Submitted: 2026-09-06
Updated: 2026-09-22
Comments: 9 pages, accepted to EMNLP 2026
Code: https://github.com/PleaseTakemeAway/MARBO
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Social deduction games (SDGs) require agents to reason under partial observability by maintaining relational beliefs about hidden roles and team alignments.
Terminology
Abstract
Social deduction games (SDGs) require agents to reason under partial observability by maintaining relational beliefs about hidden roles and team alignments. While recent LLM-agent approaches improve gameplay through prompting and preference optimization, they often optimize actions and in-game speech without explicitly grounding them in such beliefs. This frequently leads to strategically inconsistent behavior, especially for compact LLM agents. We introduce Multi-Agent Relational Belief Optimization (MARBO), a belief-grounded preference optimization framework that leverages relational beliefs to guide strategic decisions and in-game speech. MARBO provides preference feedback only when behaviors are supported by reliable relational beliefs and lead to strategically favorable social outcomes, encouraging more consistent learning under uncertainty. Experiments on representative SDGs show that MARBO enables compact LLM agents to consistently outperform existing baselines. The Code is available on https://github.com/PleaseTakemeAway/MARBO.
Sources
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction
- Training Verifiers to Solve Math Word Problems
- One Model, All Roles: Multi-Turn, Multi-Agent Self-Play Reinforcement Learning for Conversational Social Intelligence
- Shaping Zero-Shot Coordination via State Blocking
- SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Learning Strategic Language Agents in the Werewolf Game with Iterative Latent Space Policy Optimization
- ReAct: Synergizing Reasoning and Acting in Language Models
- Multi-agent KTO: Reinforcing Strategic Interactions of Large Language Model in Language Game
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Enhance Reasoning for Large Language Models in the Game Werewolf
- Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf
- RRHF: Rank Responses to Align Language Models with Human Feedback without tears
- SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection