What Does Multi-Agent LLM Debate Actually Change? A Layered Analysis of Disagreement and Answer Quality
cs.AI
Submitted: 2026-09-07
Updated: 2026-09-22
Comments: Code: https://github.com/chenmoneygithub/llm-committee
Code: https://github.com/chenmoneygithub/llm-committee
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multi-agent debate, in which several LLMs exchange arguments before answering, is widely assumed to improve answer quality by surfacing genuine disagreement.
Terminology
Abstract
Multi-agent debate, in which several LLMs exchange arguments before answering, is widely assumed to improve answer quality by surfacing genuine disagreement. That mechanism is rarely checked. We introduce four measurements: (A) the agreement a debater reports; (B) whether its reply text actually pushes back; (C) whether the position persists once the eliciting instruction is removed; and (D) for open-weight models, the stance response in the debater's own token log-probabilities. We evaluate three-model committees debating open-ended GlobalOpinionQA across 750 debates under three tones: friendly (seek common ground), neutral, and hostile (stress-test every position). (A) Tone strongly reshapes reported agreement: full agreement differs by 50.4 percentage points between the friendly and hostile endpoints. (B) A judge that reads only the reply text, never the self-report or the condition, recovers the same pattern. (C) The dissent appears partly tied to the instruction that elicited it: labels revert toward agreement 23.1 points more often after deleting the hostile instruction than under a matched re-ask that keeps it; question-weighted inference is inconclusive on first-round turns alone (p=0.0625), significant pooling all rounds (p=0.016), and only 11/28 first-round reversions also appear in the reply text. (D) Opposing arguments weaken a debater's stance margin more consistently than they shift its direction. For final answers we detect no quality gain: a bias-checked jury returns 299/299 ties (ruling out only large differences), accuracy on a verifiable control task is unchanged, and a jury without the bias check had declared debate the winner 66% of the time -- an artifact of reading order. Taken together, LLM debate readily changes what agents say, but we find much weaker evidence that it changes what they persistently endorse or improves the quality of the final answer.
Sources
- LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
- Stay Focused: Problem Drift in Multi-Agent Debate
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
- When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models
- Improving Factuality and Reasoning in Language Models through Multiagent Debate
- Towards Measuring the Representation of Subjective Global Opinions in Language Models
- LiveBrowseComp: Are Search Agents Searching, or Just Verifying What They Already Know?
- SycEval: Evaluating LLM Sycophancy
- It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty
- Not All Flips Are Conformity: Decomposing Stance Convergence in Multi-Agent LLM Debate
- AI safety via debate
- Language Models (Mostly) Know What They Know
- Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems
- On scalable oversight with weak LLMs judging strong LLMs
- The Confident Liar: Diagnosing Multi-Agent Debate with Log-Probabilities and LLM-as-Judge
- Debating with More Persuasive LLMs Leads to More Truthful Answers
- DEBATE: Devil's Advocate-Based Assessment and Text Evaluation
- Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation
- Measuring Faithfulness in Chain-of-Thought Reasoning
- More Agents Is All You Need
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection