Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate
cs.AI
Submitted: 2025-10-11
Updated: 2026-08-27
Comments: Accepted to COLM 2026
Code: https://github.com/dlab-projects/interaction
License: http://creativecommons.org/licenses/by/4.0/
The gist: As agentic AI systems are deployed in advisory and evaluative roles, understanding how multi-agent interactions shape behavior becomes essential.
Terminology
Abstract
As agentic AI systems are deployed in advisory and evaluative roles, understanding how multi-agent interactions shape behavior becomes essential. Multi-agent debate has been studied as a mechanism to improve accuracy, but less is known about how debate structure -- the interaction protocol -- affects the values, dynamics, and consensus patterns that emerge when models navigate contested, real-world decisions. We address this gap by facilitating multi-agent debates among three models (GPT-4.1, Claude 3.7 Sonnet, and Gemini 2.0 Flash) to collectively assign blame in 1,000 everyday dilemmas from Reddit's ``Am I the Asshole'' community. We compare synchronous (parallel) and round-robin (sequential) interaction protocols, mirroring two fundamental ways multi-agent systems are orchestrated in practice. Across more than 30,000 total debates, our findings show striking behavioral differences, which we characterize through two dynamics: inertia and conformity. In the synchronous setting, GPT-4.1 showed stronger inertia (0.6-3.1% revision rates) than either Claude 3.7 Sonnet or Gemini 2.0 Flash (28-41% revision rates). Meanwhile, in round-robin debates, GPT-4.1 and Gemini 2.0 Flash stood out as highly conforming relative to Claude 3.7 Sonnet, with their verdict behavior strongly shaped by order effects. We further characterized the values invoked during debate, finding that GPT-4.1 emphasized personal autonomy and honest communication relative to its debate partners, while Claude 3.7 Sonnet and Gemini 2.0 Flash prioritized empathetic dialogue. Together, these results show how interaction protocol shapes moral reasoning in multi-turn debates, establishing it as a substantive sociotechnical design consideration in multi-agent systems.
Sources
- The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test Awareness
- Moral Foundations of Large Language Models
- Large Language Models Reflect the Ideology of their Creators
- Humans or LLMs as the Judge? A Study on Judgement Biases
- ELEPHANT: Measuring and understanding social sycophancy in LLMs
- DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life
- Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
- DEBATE: A Large-Scale Benchmark for Evaluating Opinion Dynamics in Role-Playing LLM Agents
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- ProtocolBench: Which LLM MultiAgent Protocol to Choose?
- Improving Factuality and Reasoning in Language Models through Multiagent Debate
- The Moral Turing Test: Evaluating Human-LLM Alignment in Moral Decision-Making
- Large Language Models in Mental Health Care: a Scoping Review
- AI safety via debate
- Debating with More Persuasive LLMs Leads to More Truthful Answers
- MentalAgora: A Gateway to Advanced Personalized Care in Mental Health through Multi-Agent Debating and Attribute Control
- Synthetic Socratic Debates: Examining Persona Effects on Moral Decision and Persuasion Dynamics
- Large Language Models Often Know When They Are Being Evaluated
- Probing and Steering Evaluation Awareness of Language Models
- When Two LLMs Debate, Both Think They'll Win
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection