Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments

arXiv:2608.09128 · cs.CL, cs.AI, cs.MA · Submitted 2026-08-10 · Read on arXiv

Keyu He, Xuhui Zhou, Maarten Sap

Carnegie Mellon University

cs.CL, cs.AI, cs.MA

Submitted: 2026-08-10

Updated: 2026-08-11

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: Social Gym and SPA RTAN: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments Abstract LLM agents are increasingly deployed in multi-agent social settings where they must

Terminology

Summary

Social Gym and SPA RTAN: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments

Abstract

LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents. Measuring and improving these social skills is hard because, unlike math or logic, social interaction offers no objective ground truth: evaluations fall back on LLM judges, which are costly, subjective, and noisy, and models get no reliable signal to learn from. To address both, the paper introduces Social Gym, an environment of 21 multi-agent social games (e.g., Werewolves, Resistance, Spyfall) whose rule-decided outcomes make agent performance verifiable and objective, with an Elo tournament that produces a cross-game leaderboard. Benchmarking experiments show that while GPT-5-mini tops the leaderboard, no model excels at all games uniformly or in all game roles, pointing to limitations of social reasoning. Motivated by this, the paper additionally proposes SPA RTAN (Self-Play and Reflect-Transfer), a training-free self-improvement loop: a model plays a game, reflects on its trajectories and their outcomes to produce a transferable playbook, and applies that playbook in subsequent games. Results show that SPA RTAN playbooks help GPT-5-mini agents level their performance on weaker roles, but largely do not improve Qwen3-32B’s performance. Together, Social Gym and SPA RTAN offer a reproducible, verifiable foundation for measuring and improving LLM social reasoning without weight updates.

Introduction

LLM-based agents are increasingly deployed and studied as social agents in multi-party interactions with humans or other agents: interacting with humans as group mediators and companions, collaborating on shared tasks through inter-agent dialogue, role-playing as autonomous inhabitants of social sandboxes, and acting as persuaders or negotiators against other models. These settings require social interaction capabilities that go beyond single-turn question answering; agents must navigate information asymmetry, deception-utility tradeoffs, and negotiations, all while sustaining coherent role-play.

Existing evaluations of LLM social reasoning suffer from three core limitations. First, many evaluations rely on static benchmarks such as theory-of-mind questionnaires, which produce reproducible scores but cannot test sustained multi-turn behavior. Second, more recent open-ended interactive evaluations assess multi-turn interaction skills but rely on LLM-as-judge scoring, which suffers from position, verbosity, and self-enhancement biases and is inherently variable and subjective. Finally, recent works evaluate social interaction skills via verifiable rewards, but each targets a single game or domain in isolation, each covering only one aspect of social intelligence. A gap remains for evaluation that is simultaneously multi-turn, broad in domain coverage, and verifiable, i.e., determined by interaction rules rather than by subjective judgment.

To bridge this gap, the paper introduces Social Gym, an environment of 21 multi-agent social games (Werewolves, Resistance, Spyfall, Prisoner’s Dilemma, and others) spanning competitive, cooperative, and mixed-motive structures. Games like these offer a richer testbed than single-turn probes because they require sustained role-playing, coalition management, deception, and strategic information control across many turns and interactions. The unified Elo tournament that produces per-game win-rate tables and a cross-game leaderboard delivers a verifiable, multi-turn measure of LLM performance across the full breadth of social-game categories. The paper also proposes SPA RTAN, a training-free self-improvement loop, answering the question: can LLMs improve at social play without parameter updates? In SPA RTAN, a model (i) plays self-play games, (ii) reads its own trajectories and writes a transferable strategic playbook, and (iii) injects this playbook into its system prompt for subsequent games. SPA RTAN is analogous to in-context fine-tuning, but the “training data” is the model’s own gameplay and the “learned weights” are natural-language rules.

Related Work

LLM evaluations of social intelligence split into three regimes that each cover a complementary part of the space but together leave a gap. Static probes such as FANToM, ToMi, and BigToM produce reproducible question-answering scores against a fixed ground truth, but reduce social reasoning to single-turn comprehension and cannot exercise sustained behavior. Open-ended interactive evaluations, of which SOTOPIA and its SOTOPIA-π fine-tuning extension are representative, preserve multi-turn dynamics but score outcomes with LLM judges or human raters; LLM judges in particular carry documented position, verbosity, and self-enhancement biases. Single-game LLM studies on Werewolf and Avalon inherit verifiable game outcomes but each isolate one social dynamic; broader text-game environments host many games but emphasize competitive strategy over the breadth of social-cognitive demands. Social Gym fills this gap with a multi-turn, rule-decided evaluation across 21 games organized along social-cognitive axes.

Self-improvement methods for LLMs broadly divide into two families. Prompt-only approaches have agents inspect their own outputs and revise, with Reflexion and Self-Refine as the canonical references, and have since been extended to skill-library construction, where an agent accumulates transferable natural-language strategies across episodes (Voyager, ExpeL). Weight-update approaches such as SPIRAL instead use multi-agent self-play with reinforcement learning to incentivize reasoning. SPA RTAN sits in the prompt-only family and is closest in spirit to the skill-library line, but prior work in that line accumulates skills within a single domain (Minecraft for Voyager, individual reasoning tasks for ExpeL); no prior work studies whether such playbooks transfer across games, which the iterated-rounds, multi-source, and strong-to-weak experiments in Section 5 investigate.

LLM agents in multi-player games have historically been studied one game at a time. Bakhtin et al. achieve human-level play in Diplomacy with CICERO by coupling a language model to a strategic planner, demonstrating that strong play in a complex social game is possible but at the cost of heavy game-specific scaffolding. Akata et al. study LLM behavior on iterated 2×2 matrix games (Prisoner’s Dilemma, Battle of the Sexes, Stag Hunt, Chicken), confined to canonical normal-form structures. Park et al. and Hagendorff probe individual behavioral skills, such as coherent role-playing and emergent deception, in isolated sandbox or single-task settings. The common limitation is breadth: each existing line covers a single game, a single equilibrium class, or a single skill, leaving open how the same model performs across the social-cognitive spectrum. Social Gym aggregates 21 games into one environment, and SPA RTAN is tested for cross-game transfer.

Social Gym Benchmark

Social Gym is designed around two principles: (i) every game must have an algorithmically verifiable outcome (win/loss, score, survival), providing the unambiguous reward signal needed for both leaderboards and downstream RLVR training; and (ii) games must span the breadth of social intelligence, from atomic strategic primitives to long-horizon group deception, to expose distinct failure modes.

System Architecture

Social Gym extends the SOTOPIA environment loop to support arbitrary N-agent interactions. A flexible Finite State Machine (FSM) engine handles complex phase transitions (e.g., Night → Day in Werewolves; Discussion → Mission Vote → Mission Execute in Resistance). Discrete state transitions also enable downstream RL value-function estimation. A visibility layer filters every message under one of three scopes: (i) Public (all alive agents, e.g., day discussion, vote results), (ii) Team-Private (faction members only, e.g., Werewolves see each other’s night-phase kill votes), or (iii) Private (single agent, e.g., Seer inspections, role cards). This tests agents’ Theory-of-Mind reasoning, as they can only infer hidden states from permitted observations. Each game implements a uniform interface (state, available actions, visible messages, reward function), so adding a new game requires only the game-specific FSM and reward, not changes to the core engine. Twenty-one games are implemented as extensions of this shared engine.

Game Suite

The 21 games are organized along two orthogonal axes: information structure (complete information / hidden state / hidden roles) and communication mode (none / structured / free-form). This produces five categories:

  • Normal-Form Games (6): Iterated matrix games with complete information and no communication. Games: Prisoner’s Dilemma, Chicken, Battle of the Sexes, Stag Hunt, Minority Game, Rock-Paper-Scissors.

  • Economic Games (3): Multi-round resource-allocation games requiring strategic reasoning and (optionally) negotiation. Games: Public Goods Game, Centipede, Bargaining.

  • Bluffing Games (4): Hidden-state games (no fixed factions) where agents must misrepresent or correctly infer private information. Games: Liar’s Dice, Skull, Coup, Sheriff of Nottingham.

  • Hidden-Role Deduction (6): Games with hidden role assignments where agents must identify allies and enemies through dialogue. Games: Chameleon, Insider, Spyfall, Undercover, Resistance, Werewolves.

  • Social Strategy (2): Complete-information games where outcomes depend on alliance formation, persuasion, and reputation rather than hidden information. Games: Survivor, Dead Last.

Measuring success via Elo Tournament

The paper reports Elo-scale ratings estimated by a regularized Bradley–Terry maximum-likelihood fit, following the LMSYS Chatbot Arena methodology. Tournament rosters are generated by enumerating all model combinations per game and running a fixed number of episodes per combination, with role and seat assignments balanced across episodes. Within each completed episode, pairwise outcomes are extracted from the final score vector and aggregated into per-pair win/tie counts, skipping same-model pairs. In free-for-all games, every cross-model agent pair contributes one outcome, with the higher-scoring agent counted as the winner. In 2-team games, only cross-team pairs contribute outcomes, using each team’s shared score. Elo is fit on the 17 competitive games; the four cooperative games (Stag Hunt, Public Goods, Battle of the Sexes, Centipede) have no well-defined ranking and are scored by win rate instead. To capture asymmetric role performance in hidden-role games, the paper additionally reports Elo-Main (majority/cooperative role: Villager, Civilian, Non-Spy) and Elo-Alt (minority/deceptive role: Werewolf, Spy, Insider, Chameleon, Undercover).

Leaderboard Results

Seven models are benchmarked spanning closed- and open-weights access, three model families, and within-family scale: GPT-5-mini, GPT-4o, GPT-4o-mini, Qwen3-32B, Qwen3-4B, Qwen2.5-3B, and Gemma3-27B. The overall leaderboard roughly tracks general capability rankings of these models. GPT-5-mini, the newest and strongest model in the slate, tops the leaderboard at 1110 Elo; Qwen2.5-3B, the smallest (3B parameters) and oldest open checkpoint, places last at 926.

However, while overall Elo and win rate track general capability, the per-game Elos show that no model is uniformly strong: per-game rankings invert sharply. The starkest case is Qwen3-32B, first on Chicken (1328) yet last on Werewolves (817); more broadly, top-ranked models have games where they fall below the 1000 anchor or behind much weaker peers, while the smallest model (Qwen2.5-3B) places near the top on others. These inversions reflect that different games reward fundamentally different behaviors, so a single scalar masks where each model actually succeeds. This motivates the need to examine per-game model scores rather than a single scalar. The paper further disaggregates the leaderboard into per-category and per-skill capability profiles, showing that relative model strengths shift across categories: no model leads in all of them, and games involving hidden-role deduction draw the sharpest capability distinctions among models.

Role-conditioned analysis. For hidden-role games, separate Elos are examined for the minority/deceptive role (Elo-Alt) and the majority/cooperative role (Elo-Main). Across the six hidden-role games, with opponents drawn from the whole model slate, no model plays both sides at the same level, and the gap usually favors the minority/deceptive role. For the strongest models this reverses under matched-capability self-play, where the minority side is the weaker one.

Qualitative observation: parroting effect in small models. Inspecting trajectories, Qwen2.5-3B frequently parrots, i.e., paraphrases the previous speaker rather than producing an independent argument, which likely contributes to its last-place Overall Elo (926): agreeing with whoever spoke last is a near-zero-information move that gives the deceptive side cover.

SPA RTAN: Self-Play and Reflect-Transfer

The leaderboard shows that no single model dominates Social Gym: every top-ranked model has games where it underperforms peers, yet per-game ranks may invert. This motivates the research question: can a model close its own per-game gaps without weight updates, by inspecting its own gameplay and extracting reusable strategies? SPA RTAN is a simple training-free self-improvement loop with three stages:

  1. Play: The model M plays N self-play games of game G, producing trajectories τ1,..., τN.

  2. Reflect: The model is shown its own trajectories along with the final outcomes (win/loss per role) and asked to write a first-person strategic playbook covering deception, detection, persuasion, information management, coalition dynamics, and timing. The model is instructed to keep the playbook game-agnostic (no references to specific game numbers). This is denoted as R1 = Reflect(M, τ1:N).

  3. Transfer: The reflection is prepended to an agent’s system prompt for subsequent games.

Application setups test the effectiveness of SPA RTAN using the strongest LLM (GPT-5-mini) and examine four different evaluation setups: (i) within-game iterated reflection, (ii) one-source → many-target cross-game transfer, (iii) many-source → one-target multigame transfer, and (iv) strong-to-weak distillation into 6 student models on Resistance.

SPA RTAN Experiments and Results

SPA RTAN is evaluated on the asymmetric hidden-role games, asking whether iterated self-play reflection can improve a model’s win rates without parameter updates. The convention uses alt for the minority/deceptive role of a game (e.g., the Werewolves team in Werewolves or the Spies in Resistance) and main for the majority/cooperative role. In vanilla GPT-5-mini self-play the alt role’s win rate is consistently lower than the main role’s on Werewolves, Spyfall, Undercover, and Resistance, establishing an imbalance in the vanilla baseline that motivates testing whether SPA RTAN can lift the weaker side. Unless noted, every condition uses n=30 games, giving binomial 95% CIs of ≈ ±18 pp.

Same-model reflection (GPT-5-mini)

Within-game iterated R1–R4 is evaluated on three hidden-role deduction games: Werewolves, Spyfall, and Resistance, chosen because both sides have headroom in the vanilla baseline and the alt side is the structurally weaker one (alt baselines 23%, 20%, 30% respectively). GPT-5-mini generates an iterated chain of self-reflection playbooks R1, R2, R3, R4, where Rn = Reflect(M, τ1:N Rn−1) distills N=30 self-play games played under the previous round’s playbook (R0 denotes vanilla). Each Rn is injected on either the alt side or the main side, with vanilla GPT-5-mini on the other; 30 games per condition.

Result: Across rounds, the R-armed side trades win rate with the vanilla side: averaged over the three games, the alt side rises from 24% baseline to 36, 36, 43, 35% under R1–4 (peak at R3), while the main side falls from 76% to 66, 59, 67, 63% (worst at R2). Per-game peaks are non-monotonic and game-specific: Werewolves and Resistance peak on the alt side at R3 (63%, 40%), Spyfall at R2 (46%).

Interpretation: Self-reflection raises the weaker side and drops the stronger side; the optimal number of reflection rounds varies by game. Contrary to iterated-reflection methods that assume more rounds yield more gain, the bulk of the gain arrives at R1 (24% → 36%), and additional rounds redistribute rather than accumulate.

Cross-game transfer (1 → n). Werewolves, Spyfall, Chameleon, Undercover, and Resistance are used as both source and target games. For each game as source, the source’s R1 playbook is injected into one side of each of the other four as target, vs. vanilla GPT-5-mini. The two heatmaps are sign-flipped: excluding the saturated Chameleon target column, alt-side injection skews positive (median +7, max +27) and main-side injection skews negative (median −7, four cells below −20). The cross-game pattern matches the within-game finding: R1 helps the disadvantaged side and either has no effect or actively hurts the advantaged side. The match is striking because the source playbook was generated on a different game, so any useful content is not target-specific.

Multigame transfer (n → 1). For K source games, the multigame playbook is Rmulti = Reflect(M, τ1:N G1,..., τ1:N GK). Three multigame playbooks of increasing breadth are constructed: Rwcs (Werewolves + Chameleon + Spyfall), Rwcsu (+ Undercover), Rwcsur (+ Resistance). Each is evaluated on its in-distribution targets and one held-out target (except Rwcsur, whose five source games leave no held-out target). Results show that multigame does not stack: on three of the four non-saturated targets (Werewolves, Spyfall, Undercover), broader source sets either do not beat the target’s own Single-R1 or degrade it; Chameleon stays pinned at the 100% alt-side ceiling under every condition; the only exception is held-out Resistance, where Rwcsu lifts the alt side from 23% (Single-R1) to 47%. Interpretation: Learning from multiple training games does not improve transfer to a new game beyond what a single related training game already provides; the only exception, Resistance, is also the held-out game with the most baseline headroom on the alt side, consistent with the lift-the-weaker-side pattern.

Cross-model distillation

In the distillation setting, a playbook generated by a strong model Mstrong is injected into a weaker student model Mweak. The held-out multigame playbook Rwcsu (trained on Werewolves + Chameleon + Spyfall + Undercover by Mstrong = GPT-5-mini, Resistance excluded) is tested when injected into weaker student models. Six students (Qwen3-32B, Qwen3-4B, Qwen2.5-3B-Instruct, Gemma3-27B, GPT-4o, GPT-4o-mini) each play one side of Resistance against vanilla GPT-5-mini, with and without the playbook.

Result: Distillation reproduces the same side asymmetry across both transfer axes. On the disadvantaged Spies side, three students gain +13 to +24 pp from the shared Rwcsu playbook; on the favored Resistance side, every student moves by at most ±5 pp. Two students (Qwen2.5-3B, Gemma3-27B) do not gain on either side. Interpretation: The same playbook produces side-dependent rather than student-dependent effects: it lifts whichever student is playing the structurally weaker role and leaves the other alone. Combined with the multigame result, the same Rwcsu playbook now lifts the underperforming side across three setups (within-model held-out target, across-model held-out target, both at once), all without ever having seen Resistance during reflection.

Open-weights model replication (Qwen3-32B)

The patterns from Sections 5.1 and 5.2 are tested on an open-weights model, Qwen3-32B, mirroring the within-model setup and cross-model setup.

Self-reflection (Open-weights model). Following the within-game protocol, Qwen3-32B runs as the only model (reflection generator, training self-play, and evaluation opponent) across six games: Werewolves, Spyfall, Resistance, Chameleon, Undercover, and Prisoner’s Dilemma. PD is included as a pure action-channel game (no chat phase) for contrast with the discussion-heavy games. Iterated playbooks R1–R4 are generated from Qwen3-32B self-play and the R-armed-side win rate per round is reported.

Result: Only Prisoner’s Dilemma shows a clean R-armed-side lift (13% → 58% at R1, +45pp). Resistance shows a smaller positive effect (+34pp on the Resistance side). The other four games (Werewolves, Chameleon, Spyfall, Undercover) are flat across all four iterated rounds and across both within-game and cross-game source playbooks.

Interpretation: The open-weights replication is largely a null result, with PD as the only clean exception. This is attributed to model capacity: at 32B parameters Qwen3-32B appears incapable of learning from the trajectories of long, complex social-deduction games, except when the required action collapses to a single discrete token (PD’s defect, Resistance’s private succeed/fail vote). One concrete symptom is that Qwen3-32B’s discussion-phase outputs frequently parrot or paraphrase the immediately prior speaker rather than producing independent content. The model is in fact self-aware enough to identify this behavior, and iterated reflection codifies it into the playbook itself, but the model does not eliminate the parroting in subsequent play.

Distillation (Open-weights model). The distillation setup is tested with Qwen3-32B as the teacher. The source playbooks are Qwen3-32B’s within-game R1 for PD and its held-out multigame Rwcsu for Resistance (the two games where Qwen3-32B’s own self-reflection produced a measurable lift). Transfer is tested to Qwen3-4B, Qwen2.5-3B, and Gemma3-27B against vanilla Qwen3-32B (30 games per side per condition). Result: PD shows clean positive distillation across all three students (+17 to +77pp with the reflection applied); Resistance is mixed (Qwen2.5-3B +10pp, Qwen3-4B −20pp, Gemma3-27B −10pp on the Resistance side). Distillation reinforces the finding: the action-channel game (PD) transfers cleanly to smaller students, while the discussion-heavy game (Resistance) does not.

Conclusion and Discussion

The paper presented Social Gym, an environment of 21 multi-agent social games organized into five categories (normal-form, economic, bluffing, hidden-role deduction, social strategy), with a unified Elo leaderboard that reveals large per-game ranking inversions and role-conditioned imbalances beneath an overall ranking that tracks general capability. The paper then introduced SPA RTAN, a training-free self-improvement loop, and evaluated it across within-model iteration, cross-game transfer, and cross-model distillation. A single regularity emerges from all three perspectives: the playbook lifts the structurally weaker side of an asymmetric game.

Two main implications are drawn. First, Social Gym and SPA RTAN jointly show that LLM social ability is not a single scalar capability: model rankings, role advantages, and reflection gains all depend strongly on the interaction structure of the game. By combining a broad game suite with targeted playbook interventions, structural properties of a social setting can be separated from model-specific failures such as weak deception, poor coalition tracking, or parroting behavior. Second, Social Gym provides a natural testbed for future work to explore reinforcement learning with verifiable rewards (RLVR): every episode produces an objective, rule-computed outcome while still requiring rich language-based interaction. This would enable future research to train and evaluate social reasoning skills at scale without LLM judges, while also testing whether learned strategies transfer across cooperation, negotiation, bluffing, and hidden-role deduction games.

Future work includes testing whether playbooks learned in rule-based games carry into realistic deployment settings, such as professional negotiation, customer-service de-escalation, or collaborative multi-agent work, extending SPA RTAN’s transfer evaluation from held-out games to held-out domains.

Limitations

External validity: All games in Social Gym have fixed rules, fixed role structures, and rule-decided outcomes. This is what lets every episode be scored without an LLM judge, but many real social interactions have no clear win/loss criterion and cannot be reduced to a single score. Whether the capabilities and playbooks measured in these games carry over to such settings remains open.

Sample size per condition is modest: 30 evaluation games per condition, giving binomial 95% CIs of roughly ±18 pp for a single-coin observation. Several effects reported are within this range and should be replicated at larger sample sizes.

No placebo-playbook control: R-armed players are compared against vanilla opponents but not against opponents armed with a content-matched placebo (scrambled or unrelated text of equal length). Without this control, playbook-content effects cannot be fully disentangled from generic prompt-perturbation effects, though the structural patterns reported (action-channel vs. free-discussion games) and the Undercover monotonic regression toward 0% argue against a pure prompt-perturbation reading.

The reflection is constrained to natural-language prose: SPA RTAN does not allow the model to update tools, retrieve external knowledge, or perform structured reasoning beyond what fits in the system-prompt text. Methods that combine reflection with retrieval or scratchpads may exhibit qualitatively different transfer behavior.

Ethics / Broader Impacts

Social Gym and SPA RTAN measure and, in some settings, improve capabilities: deception, persuasion, coalition manipulation. Thus, they may carry dual-use risk if transferred from games to real interactions involving humans. Three mitigating factors are noted: all experiments are confined to fully synthetic multi-agent games with no human subjects; the improvements are training-free, modest in size, and largely null for open-weights models; and the verifiable-reward framing is intended primarily as an evaluation tool for diagnosing such capabilities rather than a recipe for deploying manipulative agents. The code is released to support reproducible measurement of these behaviors, and use of the playbook-distillation procedure in adversarial human-facing applications is discouraged.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems:

Improvement: Implement a dynamic strategy selection mechanism that detects whether the AI occupies a structurally disadvantaged role in asymmetric interactions and automatically adjusts its behavior to compensate.

What the improved AI can do:

  • Recognize when it's in a minority/deceptive position (e.g., spy, werewolf, undercover agent) versus a majority/cooperative position

  • Automatically generate and apply targeted strategies for the weaker role rather than using a one-size-fits-all approach

  • Shift from passive agreement/parroting to independent argumentation when in disadvantaged positions, addressing the specific failure mode observed in smaller models

Improvement: Build a reflection system that doesn't assume more reflection rounds equal better performance, but instead detects when additional reflection stops providing gains and halts the process.

Improvement: Create a transfer learning system that extracts game-agnostic strategic principles and applies them to new domains, but only injects them into the side that needs improvement.

Improvement: Implement a capability-gating mechanism that determines whether the AI can meaningfully learn from complex multi-turn trajectories, and if not, falls back to simpler action-channel learning.

Improvement: Use the rule-decided outcomes from multi-agent games as objective training signals, replacing subjective LLM-judge evaluations.

Improvement: Build a knowledge distillation system where strong models generate strategic playbooks that are injected into weaker models, with the understanding that effects are side-dependent rather than student-dependent.

Improvement: Implement a system that categorizes interactions by their structural properties (information asymmetry, communication mode, role structure) and selects appropriate strategies based on these categories rather than treating all social interactions the same.

Abstract

LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents. Measuring and improving these social skills is hard because, unlike math or logic, social interaction offers no objective ground truth: evaluations fall back on LLM judges, which are costly, subjective, and noisy, and models get no reliable signal to learn from. To address both, we first introduce Social Gym, an environment of 21 multi-agent social games (e.g., Werewolves, Resistance, Spyfall) whose rule-decided outcomes make agent performance verifiable and objective, with an Elo tournament that produces a cross-game leaderboard. Benchmarking experiments show that while GPT-5-mini tops the leaderboard, no model excels at all games uniformly or in all game roles, pointing to limitations of social reasoning. Motivated by this, we additionally propose SPaRTan (Self-Play and Reflect-Transfer), a training-free self-improvement loop: a model plays a game, reflects on its trajectories and their outcomes to produce a transferable playbook, and applies that playbook in subsequent games. Our results show that SPaRTan playbooks help GPT-5-mini agents level their performance on weaker roles, but largely do not improve Qwen3-32B's performance. Together, Social Gym and SPaRTan offer a reproducible, verifiable foundation for measuring and improving LLM social reasoning without weight updates.

Sources

Related papers