Mechanism Design Is Not Enough: Prosocial Agents for Cooperative AI
cs.GT, cs.AI
Submitted: 2026-05-08
Updated: 2026-09-22
Comments: 42 pages. Accepted at ICML 2026 AIWILD
Code: https://github.com/Flecart/prosocial-agents
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Ensuring that AI agents behave safely and beneficially when interacting with other parties has emerged as one of the central challenges of modern AI safety.
Terminology
Abstract
Ensuring that AI agents behave safely and beneficially when interacting with other parties has emerged as one of the central challenges of modern AI safety. While mechanism design, as the theory of designing rules to align individual and collective objectives, can incentivize cooperative behavior, it is still an open question whether it alone is sufficient to maximize LLM agents' social welfare. This work proves that the answer is negative: drawing from incomplete contract theory, we formally show that when contracts cannot distinguish all relevant future contingencies, there is a strictly positive welfare loss that no realistic mechanism can eliminate. We show that prosocial agents, who weigh others' welfare alongside their own, can close this gap and achieve outcomes that are socially superior and individually beneficial. Experimentally, we show that in multi-agent resource-allocation environments and canonical social dilemmas where agents are powered by large language models, prosociality is beneficial. The implication for AI safety is clear: to enable cooperative interactions at scale, designing adequate mechanisms is not sufficient; agents must be built to be intrinsically prosocial.
Sources
- Self-Resource Allocation in Multi-Agent LLM Systems
- Steering Large Language Model Activations in Sparse Spaces
- GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory
- Cooperative and uncooperative institution designs: Surprises and problems in open-source game theory
- Evaluating Cooperation in LLM Social Groups through Elected Leadership
- An Interpretable Automated Mechanism Design Framework with Large Language Models
- Code Simulation as a Proxy for High-order Tasks in Large Language Models
- Large Language Models Often Know When They Are Being Evaluated
- GPT-4o System Card
- Learning Robust Social Strategies with Large Language Models
- Moral Alignment for LLM Agents
- CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas
- GemNet: Menu-Based, Strategy-Proof Multi-Bidder Auctions Through Deep Learning
- Understanding the Mechanism of Altruism in Large Language Models
Related papers
- Exact Regret Frontiers and Externality Scheduling in Centralized Serial-Dictatorship Bandits
- In-Context Credit Assignment via the Core
- Breaking 1/epsilon Barrier in Quantum Zero-Sum Games: Generalizing Metric Subregularity for Spectraplexes
- Enhancing Affine Maximizer Auctions with Correlation-Aware Payment
- LLM Bidders Preserve the Mechanism-Level Orderings of Human Bidders
- Towards Performatively Stable Equilibria in Decision-Dependent Games for Arbitrary Data Distribution Maps