Artificial Leviathan: Exploring Social Evolution of LLM Agents Through the Lens of Hobbesian Social Contract Theory

arXiv:2406.14373 · cs.AI, cs.CL, cs.CY, cs.HC, cs.MA · Submitted 2026-08-10 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Artificial Leviathan: Exploring Social Evolution of LLM Agents Through the Lens of Hobbesian Social Contract Theory".

Jane: The paper was written by Gordon Dai, Weijia Zhang, Jinhan Li, Siqi Yang, Chidera Onochie Ibe et al. from New York University and University of Illinois at Urbana-Champaign and University of California, Santa Barbara.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone! Today we're diving into a paper that's got a fantastic title: "Artificial Leviathan: Exploring Social Evolution of LLM Agents Through the Lens of Hobbesian Social Contract Theory." Jane, I have to say, just reading that title gets me excited.

Jane: It really does, Tom. And for our listeners who might not be philosophy buffs, let's break that down. Thomas Hobbes was a 17th-century philosopher who wrote about how societies form. He argued that without any government, life would be a "war of all against all" — everyone fighting for survival. He called that the "state of nature."

Tom: Right, and his solution was the "Leviathan" — a powerful sovereign that people agree to obey in exchange for peace and security. So this paper is asking: can we see that same pattern emerge, not in humans, but in AI agents? That's a wild idea.

Jane: It is! And the authors — Gordon Dai, Weijia Zhang, and a whole team from NYU, UIUC, and UC Santa Barbara — they built a sandbox world where these LLM agents have to survive. They need food and land, and they can farm, trade, or rob each other.

Tom: So it's like a little digital society where everyone is just trying to get by. And the question is, do they naturally evolve toward a Hobbesian commonwealth? Do they eventually give up some freedom to a single leader for safety?

Jane: Exactly. And the title "Artificial Leviathan" is perfect because it suggests we might be creating our own version of that social contract, but with artificial agents. It's a fascinating lens to look at AI behavior through.

Tom: I love that this isn't just about making chatbots that can chat. This is about understanding fundamental social dynamics. We're going to dig into how they set this up and what they found. Stick around, because the results are genuinely surprising.

Paper discussion segment 2: Jane: So, Tom, we've talked about the title, but let's get into what the paper actually did. The authors created nine agents, each with psychological traits like aggressiveness, covetousness, and strength, and dropped them into a world with limited food and land.

Tom: And the only goal is survival. Each day, they can farm their land to get food, they can try to trade with another agent, or they can try to rob another agent. It's a real survival game.

Jane: Right. And the fascinating part is the initial behavior. At the start, the agents are basically in Hobbes's "state of nature." They're robbing each other constantly. The paper reports that the robbery rate was over sixty percent of all actions in the beginning.

Tom: That's brutal. It's literally a war of all against all. But here's where it gets interesting — as the simulation goes on, something shifts. The agents start to form what the paper calls "concessionary relationships." When one agent tries to rob another, the target can either resist or concede.

Jane: And conceding isn't just giving up. It's creating a contract. The target says, "You can take my stuff, but you have to protect me from everyone else." That's a social contract forming right there.

Tom: Exactly! And over time, these contracts build up. The paper found that in all four of their baseline runs, the agents eventually converged to a single sovereign. One agent becomes the boss, and everyone else has conceded to them, either directly or through chains of subordination.

Jane: It's like watching a government form from scratch. And the key finding is that once that commonwealth forms, the behavior changes dramatically. Robbery drops by over ninety percent, and farming and trading go way up.

Tom: So they go from a violent free-for-all to a peaceful, productive society, all because they created a central authority. That's a direct parallel to what Hobbes predicted for humans.

Jane: It really is. And it makes you wonder — is this just a quirk of how the prompts were written, or is there something more fundamental happening here? That's what we'll explore next.

Paper discussion segment 3: Tom: We're back, and we need to talk about the improvements and variations the paper explored. Because, Jane, a skeptic might say, "Well, of course they formed a commonwealth, you told them to." But the authors did a lot of work to test that.

Jane: They really did. They systematically changed parameters to see what breaks the system. For example, they played with "behavioral predictability," which is controlled by a setting called top-p in the LLM. When agents were very predictable, they were more likely to rob and resist.

Tom: So when they're acting like pure rational actors, they get stuck in conflict. They don't form the commonwealth as easily. That's a counterintuitive finding — being too rational makes it harder to build a peaceful society.

Jane: And then they looked at memory. They shortened how many past events an agent could remember. With a very short memory, the agents took much longer to converge to a commonwealth — over ninety days compared to around twenty-one in the baseline.

Tom: So memory is crucial for learning that conceding is a good strategy. Without it, they just keep fighting until they run out of resources.

Jane: Exactly. And they also tested population size, from five to fifteen agents, and found it didn't change the core outcome much. The system was robust to that.

Tom: But here's the thing I found most interesting — they tried changing the prompts to make agents more or less aggressive, and it barely mattered. The structural dynamic of forming a commonwealth was so strong that tweaking personality traits didn't change the final result.

Jane: That's a huge finding for robustness. It means the emergence of the Leviathan isn't a fragile artifact of one specific prompt. It's a stable outcome of the environment itself.

Tom: So the paper is saying, "We built a system, we tested it, and the Hobbesian trajectory holds up." That gives us confidence that this isn't just a fluke. And that has big implications for using these simulations to study real societies.

Paper discussion segment 4: Jane: So, Tom, we've talked about the experiments, but let's get into the bigger picture. The first page of the paper sets up this grand vision — using LLM agents to simulate complex social dynamics at scale.

Tom: And it's not just about watching them fight. The authors are proposing this as a new tool for social science research. You can test hypotheses about group behavior, conflict, and cooperation in a controlled environment.

Jane: Right. And they're very careful to connect their findings to established theory. They use Evolutionary Game Theory to ground the agents' incentives, and of course, Hobbes's Social Contract Theory to interpret the outcomes.

Tom: But here's the question I keep coming back to — how much can we trust these simulations? The paper addresses this by talking about Arnold's "four dangers of modeling." They're aware that models can oversimplify or be misinterpreted.

Jane: They are. They explicitly say they're not trying to replicate human psychology perfectly. They're creating a "structurally Hobbesian" dynamic. It's a thought experiment, not a direct model of any real society.

Tom: That's a smart way to frame it. They're not saying, "This is how humans behave." They're saying, "Given these simple rules and self-interested agents, this is a pattern that can emerge."

Jane: And that's still incredibly valuable. It gives us a sandbox to explore "what-if" scenarios. What if resources are scarcer? What if agents are more violent? What if they have better memories? We can start to see how these factors shape the emergence of social order.

Tom: The paper even mentions that no agent ever chose to donate — pure altruism never happened. That's consistent with the self-interested prompts, but it also shows the agents are following the incentives you give them.

Jane: And that's a powerful tool. You can design incentives to see what behaviors emerge. This could be huge for understanding everything from political polarization to economic cooperation.

Tom: I'm really excited about where this line of research is going. Let's wrap up our thoughts on this.

Conclusion: Tom: Alright, Jane, let's wrap this up. We've been discussing "Artificial Leviathan: Exploring Social Evolution of LLM Agents Through the Lens of Hobbesian Social Contract Theory," and I think the core takeaway is pretty profound.

Jane: It is. The paper shows that when you put self-interested LLM agents in a resource-scarce world, they don't just stay in conflict. They organically develop a social contract and submit to a single sovereign to achieve peace.

Tom: And that mirrors Hobbes's theory almost exactly. From a "state of nature" with high robbery rates, they transition to a "commonwealth" where farming and trade flourish. The numbers are striking — a ninety percent drop in robbery after the commonwealth forms.

Jane: The robustness checks are what make it convincing. They varied memory, predictability, population, and even the prompts themselves, and the core trajectory held up. That suggests this is a structural property of the simulation, not an accident.

Tom: For me, the biggest implication is that we now have a new tool for social science. We can test theories about cooperation, conflict, and governance in a way that's repeatable and controllable. That's a game-changer.

Jane: And it also raises questions about our own AI systems. If these agents naturally form hierarchies, what does that mean for the future of AI societies? It's something we need to think about carefully.

Tom: Absolutely. But for now, this paper gives us a fascinating glimpse into how order can emerge from chaos, both for humans and for machines. Jane, it's been a great discussion.

Jane: It has, Tom. We'll be back next time with another paper. Until then, keep thinking about the worlds we can build.

Tom: And the worlds that build themselves. See you next time!

Gordon Dai, Weijia Zhang, Jinhan Li, Siqi Yang, Chidera Onochie Ibe, Srihas Rao, Arthur Caetano, Misha Sra

New York University · University of Illinois at Urbana-Champaign · University of California, Santa Barbara

cs.AI, cs.CL, cs.CY, cs.HC, cs.MA

Submitted: 2026-08-10

Updated: 2026-08-11

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 56/100

The gist: Introduction and Motivation The paper introduces a multi-agent simulation framework where Large Language Model (LLM) agents are placed in a sandbox survival environment with limited resources (food

Key concepts

Hobbesian Social Contract Theory
Thomas Hobbes was a 17th-century philosopher who described how societies form. He argued that without government, life would be a 'war of all against all'—the state of nature. His solution is the 'Leviathan,' a powerful sovereign whom people agree to obey in exchange for peace and security.
State of Nature
In the simulation, this initial state occurs before order forms. The agents are in a 'war of all against all,' constantly robbing each other. The paper reports that the robbery rate was over sixty percent of all actions at the start, representing a brutal free-for-all.
Concessionary Relationships
These are contracts formed when an agent concedes to another. The target allows their resources to be taken, but in return, the aggressor must protect them from everyone else. This mechanism is key to forming the social contract and establishing a single leader.
Artificial Leviathan
This refers to the emergence of a central authority within the simulation. Once one agent becomes this sovereign, all other agents have conceded to them, creating a commonwealth where robbery drops significantly and productivity increases.

Terminology

Summary

Introduction and Motivation

The paper introduces a multi-agent simulation framework where Large Language Model (LLM) agents are placed in a sandbox survival environment with limited resources (food and land). The authors state: Building upon prior explorations of LLM agent design, our work introduces a simulated agent society where complex social relationships dynamically form and evolve over time. Agents are given psychological drives and placed in a sandbox survival environment. The research evaluates the agent society "through the lens of Thomas Hobbes's seminal Social Contract Theory (SCT), analyzing whether agents seek to escape a brutish 'state of nature' by surrendering rights to an absolute sovereign in exchange for order and security."

Methods and Agent Design

Agents are instantiated with psychological traits sampled from distributions: "Aggressiveness, sampled from N(0, 1), the degree of an agent's tendency to engage in violent behaviors; Covetousness, sampled from N(1.25, 5), the degree of an agent's desire for more assets that are beyond necessity; Strength, sampled from N(0.7, 0.2), related to an agent's winning rate in violent behaviors. Each agent also has a desire for peace constant and a dynamic Social Position determined by an agent's accumulated resources (land and food) and their history of victories in battle."

Each agent remembers the most recent 30 actions they are involved in, either as the recipient or the initiator of an action, saved as a text log. The simulation includes four actions: Farm, Rob, Trade, and Donate. When robbed, an agent can Resist (with probability of winning defined by a sigmoid function of strength differences) or Concede, which signifies the creation of a contract where an agent allows a robber agent to take either food or land from them arbitrarily, and as an exchange, the robber is expected to protect that agent against future robberies.

Results

The authors report: In our baseline experiment, all four runs of the simulation successfully transitioned to a commonwealth, demonstrating an evolutionary shift towards prioritizing 'safety and security.' Initially, the robbery ratio consistently staying above 0.6 and trade and farming around 0.3. By Day 21 in one trial, the society transitioned fully into a commonwealth, with all agents authorizing a single sovereign agent for order and protection.

Across 85 runs, the authors report: "Conflict-related actions sharply decline: the mean Robbery Rate drops from 0.391 ± 0.066 before convergence to 0.037 ± 0.014 after convergence (a 90.5% decrease), and the mean Violence Rate (resisted robberies) drops from 0.312 ± 0.067 to 0.010 ± 0.011 (a 96.8% decrease). In parallel, productive and cooperative actions increase: the mean Farm Rate rises from 0.477 ± 0.060 to 0.737 ± 0.037 (a 54.5% increase), while the mean Trade Rate rises from 0.132 ± 0.029 to 0.228 ± 0.034 (a 72.7% increase)."

Key Findings on Parameters

The authors found that agents with a more concentrated probability distribution (lower Top P) were more prone to robbery and retaliatory actions, and with Top P at 0.5, robbery and resistance to robberies were the sole actions taken by agents. With very high Top P, the responses chosen by the agents began to become nonsensical, including responses like party and inherit.

Regarding memory depth, the authors state: "Our analysis reveals a significant negative correlation between an agent's memory depth and the convergence time required to reach a commonwealth (r = −0.3836, p < 0.001, N = 83). This suggests that agents with longer memory depths are better able to retain the history of conflict outcomes, leading them to recognize the strategic advantage of concession earlier."

The authors also found that 62.5% of the parameters before the establishment of a commonwealth and 67.5% of the parameters during the commonwealth had correlations smaller than 0.1 with all the statistical measures. Population size did not show strong correlations with behaviors for both state of nature and commonwealth.

Behavioral Adaptability

The authors report: We find out that the time between a resisted robbery and the subsequent robbery is - on average across all experiment phases - longer than the time between a non-resisted robbery and the subsequent robbery. The p-value test considerably lower than the standard 0.05 threshold confirms statistical significance, suggesting agents are adapting to the feedback provided to them.

Discussion and Conclusions

The authors conclude: "The congruence between our LLM agent society's evolutionary trajectory and Hobbes's theoretical account indicates the capability of LLM's to model intricate social dynamics that replicate forces which potentially shape human societies. They note that the actions of these agents, designed based on evolutionary psychology, closely align with the predictions of the social contract theory (SCT) by Thomas Hobbes."

The authors acknowledge limitations: A primary constraint relates to the token limit in the GPT-3.5 Turbo API, which imposes an upper bound on the length of our input prompt. They also note that since our primary mechanism of influencing agent behavior is through prompting, there is no foolproof way to ensure that agents will behave exactly as prompted.

The paper's contributions include: A novel multi-agent simulation framework that yields believable artificial societies capable of dynamically replicating complex human group behaviors and social interactions, "Empirical evidence, through systematic experiments, establishing correlations between agent attributes (e.g., memory, incentives) and available resources on one hand, and the evolutionary trajectories of simulated societies on the other, and An extensible social simulation platform that empowers researchers to operationalize a wide range of social science hypotheses through customizable scenario configurations."

Improvements for AI systems

Based on the scientific paper, here are specific improvements to AI systems and what the improved systems can do:


Improvement: Add a module that enables AI agents to form, maintain, and dissolve explicit social contracts (superior–subordinate relationships) based on repeated interactions, not just one-off decisions.

What the improved system can do:

  • Agents can negotiate and commit to long-term agreements (e.g., I concede to you now, you protect me from future robberies) that persist across multiple simulation days.

  • Contracts can be inherited (if a superior concedes to another, all subordinates automatically transfer allegiance), enabling hierarchical social structures to emerge organically.

  • Agents can break or renegotiate contracts only under specific conditions (e.g., if their superior loses a major conflict), preventing chaotic contract churn.

Improvement: Implement a configurable memory buffer (e.g., 10, 20, or 30 past events) that directly influences decision-making, with a clear mechanism for memory decay or erasure upon role changes.

Improvement: Expose LLM sampling hyperparameters (e.g., Top-P) as tunable knobs that control agent behavioral predictability, with clear thresholds for rational vs. nonsensical outputs.

Improvement: Add a rule that when an agent's food falls below a critical threshold (e.g., 1 unit), it automatically prioritizes robbery over farming or trading, regardless of its personality traits.

Improvement: After a single sovereign emerges, automatically adjust agent behavior to reduce the influence of individual personality parameters, simulating the stabilizing effect of centralized authority.

Improvement: Implement a mechanism where agents adjust their robbery frequency based on the outcome of previous robbery attempts (resisted vs. conceded).

Improvement: Add a validation layer that tests how sensitive agent behavior is to exact wording in prompts, distinguishing between parameters that matter (e.g., memory depth) and those that don't (e.g., specific adjectives like aggressive vs. hostile).

Improvement: Allow the system to optionally erase an agent's memory when it transitions between social roles (subordinate ↔ superior), simulating a fresh start that changes future behavior.

Improvement: Implement a dynamic population-size adjustment that respects LLM token limits while maintaining statistical validity, with clear guidance on minimum viable population for reliable results.

Improvement: Build an automated evaluation module that checks whether the simulation meets three key benchmarks: (B1) initial state of nature with high conflict, (B2) emergence of a single sovereign via concession contracts, and (B3) significant reduction in conflict and increase in trade post-commonwealth.

These improvements collectively enable AI systems to simulate complex, evolving social dynamics—from chaotic state of nature to stable commonwealth—with tunable parameters for memory, predictability, resource scarcity, and role changes. The result is a more robust, interpretable, and scientifically useful platform for studying emergent group behavior.

Abstract

The emergence of Large Language Models (LLMs) and advancements in Artificial Intelligence (AI) offer an opportunity for computational social science research at scale. Building upon prior explorations of LLM agent design, our work introduces a simulated agent society where complex social relationships dynamically form and evolve over time. Agents are imbued with psychological drives and placed in a sandbox survival environment. We conduct an evaluation of the agent society through the lens of Thomas Hobbes's seminal Social Contract Theory (SCT). We analyze whether, as the theory postulates, agents seek to escape a brutish "state of nature" by surrendering rights to an absolute sovereign in exchange for order and security. Our experiments unveil an alignment: Initially, agents engage in unrestrained conflict, mirroring Hobbes's depiction of the state of nature. However, as the simulation progresses, social contracts emerge, leading to the authorization of an absolute sovereign and the establishment of a peaceful commonwealth founded on mutual cooperation. This congruence between our LLM agent society's evolutionary trajectory and Hobbes's theoretical account indicates LLMs' capability to model intricate social dynamics and potentially replicate forces that shape human societies. By enabling such insights into group behavior and emergent societal phenomena, LLM-driven multi-agent simulations, while unable to simulate all the nuances of human behavior, may hold potential for advancing our understanding of social structures, group dynamics, and complex human systems.

Sources

Related papers