A latent dimension of Condorcet's jury theorem for multiple AI advisers
cs.CY, cs.AI, cs.HC, cs.MA
Submitted: 2026-09-13
Updated: 2026-09-28
Comments: 11 pages, 4 figures, 1 table
License: http://creativecommons.org/licenses/by/4.0/
The gist: When the same question is asked of multiple AI advisers, as in self-consistency and LLM-as-a-judge panels, Condorcet's jury theorem predicts that adding independent, competent advisers makes the
Terminology
Abstract
When the same question is asked of multiple AI advisers, as in self-consistency and LLM-as-a-judge panels, Condorcet's jury theorem predicts that adding independent, competent advisers makes the majority more reliable. The theorem, however, has a latent dimension when viewed from the user's vantage: adding advisers also makes disagreement more visible. A binomial model reveals that this ``visible dissent'' becomes nearly inevitable as the number of advisers grows, and that reliability and disagreement both approach certainty but at different convergence rates. The two rates cross at an adviser accuracy of 4/5 (0.8). Below this value, visible dissent approaches certainty faster than reliability and, with enough advisers, becomes more likely than a correct majority. Even ideal panels of independent and competent advisers can be correct in aggregate but appear divided; such disagreement does not by itself indicate aggregation failure. The way advisers split also provides a common basis for predictive multiplicity, reconciliation load, and reliance miscalibration. These results indicate two distinct decisions when using multiple AI advisers: how many advisers to consult and how their verdicts should be presented and interpreted.
Sources
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles
- DiscoUQ: Structured Disagreement Analysis for Uncertainty Quantification in LLM Agent Ensembles
- Increasing LLM response trustworthiness using voting ensembles
- Epistemic Filtering and Collective Hallucination: A Jury Theorem for Confidence-Calibrated Agents
- Algebraic Evaluation Theorems
- LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning
- Benchmark Illusion: Disagreement among LLMs and Its Scientific Consequences
- A Statistical Framework for Auditing Behavioral Dependence and Induced Bias in LLM Judges
- Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels
- When Is Collective Intelligence a Lottery? Multi-Agent Scaling Laws for Memetic Drift in LLMs
Related papers
- Reasoning Enhances Robustness to Prompt Injection in LLM-Based Consensus
- Generative AI Purpose-built for Social and Mental Health: A Real-World Pilot
- PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
- What is an intelligent system?
- AI University: An LLM-Powered Learning Assistant for Engineering---A Finite Element Method Case Study
- Generative AI Use in Entrepreneurship: An Integrative Review and an Empowerment-Entrapment Framework