Confidence Composition for Multiagent Language Model Systems
cs.AI, cs.LG, cs.MA
Submitted: 2026-06-11
Updated: 2026-09-21
Comments: 22 pages and 4 figures in total, 11 pages (8 main, 3 reference) and 2 figures excluding the appendix
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Multiagent language model systems, such as collaborative reasoning and debate, produce multiple correlated candidate answers and confidence signals.
Terminology
Abstract
Multiagent language model systems, such as collaborative reasoning and debate, produce multiple correlated candidate answers and confidence signals. However, these signals are usually calibrated only at the individual agent level, and provide no principled confidence estimate for the system's final answer. We formulate this as a confidence composition problem where combining confidence across agents and reasoning stages while preserving both selective utility and probabilistic reliability. We study confidence-aware routing and log-odds pooling protocols that select among candidate answers and output a system-level confidence. Across five benchmarks, 30 heterogeneous and homogeneous model pairs, and two confidence estimators, our gated-fusion methods improve AUARC and reduce Brier score over single agent, standard debate, and selective debate baselines, while retaining competitive weighted F1-score as a correctness metric. We further show that our log-odds fusion is overconfident due to correlated intermediate signals. We propose a shared dependence discount that substantially improves reliability while preserving predictions.
Sources
- Debate Only When Necessary: Adaptive Multiagent Collaboration for Efficient LLM Reasoning
- Phi-4 Technical Report
- The Llama 3 Herd of Models
- Qwen Technical Report
- Enhancing Multi-Agent Debate System Performance via Confidence Expression
- Breaking the Martingale Curse: Multi-Agent Debate via Asymmetric Cognitive Potential Energy
- Epistemic Gain, Aleatoric Cost: Uncertainty Decomposition in Multi-Agent Debate for Math Reasoning
- Gemma 2: Improving Open Language Models at a Practical Size
- Can LLM Agents Really Debate? A Controlled Study of Multi-Agent Debate in Logical Reasoning
- Talk Isn't Always Cheap: Understanding Failure Modes in Multi-Agent Debate
- Demystifying Multi-Agent Debate: The Role of Confidence and Diversity
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection