Measuring Collapse and Correction in Homogeneous-Panel LLM Debate
cs.CL
Submitted: 2026-09-28
Updated: 2026-09-28
Project page: https://lixin.ai/DebateLedger
Terminology
Sources
- CGES: Confidence-Guided Early Stopping for Efficient and Accurate Self-Consistency
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Debate Only When Necessary: Adaptive Multiagent Collaboration for Efficient LLM Reasoning
- Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems
- LLMs Know They're Wrong and Agree Anyway: The Shared Sycophancy-Lying Circuit
- The Sequential Edge: Inverse-Entropy Voting Beats Parallel Self-Consistency at Matched Compute
- The Value of Variance: Mitigating Debate Collapse in Multi-Agent Systems via Uncertainty-Driven Policy Optimization
- Simple synthetic data reduces sycophancy in large language models
- Can LLM Agents Really Debate? A Controlled Study of Multi-Agent Debate in Logical Reasoning
- Talk Isn't Always Cheap: Understanding Failure Modes in Multi-Agent Debate
- Knowledge Divergence and the Value of Debate for Scalable Oversight
- Disentangling the Drivers of LLM Social Conformity: An Uncertainty-Moderated Dual-Process Mechanism
- Demystifying Multi-Agent Debate: The Role of Confidence and Diversity
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering