When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines
cs.MA, cs.AI
Submitted: 2026-03-20
Updated: 2026-07-21
Comments: v2: Updated author list to match the published version. Published in Applied Sciences (MDPI) 2026, DOI: 10.3390/app1010000
Journal ref: Applied Sciences 16(10), 4914 (2026)
DOI: 10.3390/app16104914
Code: https://github.com/maryanskyy/agents-disagree-experiments
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multi-agent LLM pipelines produce contradictory evidence on whether team diversity improves output quality: heterogeneous Mixture-of-Agents teams outperform single models, yet homogeneous Self-MoA
Terminology
Abstract
Multi-agent LLM pipelines produce contradictory evidence on whether team diversity improves output quality: heterogeneous Mixture-of-Agents teams outperform single models, yet homogeneous Self-MoA teams consistently win under synthesis-based aggregation. We propose a resolution by identifying the selection bottleneck -- a crossover threshold in aggregation quality that determines whether diversity helps or hurts. Under this model, we obtain a closed-form crossover threshold s* (Proposition 1) that separates the regimes where diversity helps and hurts. In a targeted experiment spanning 42 tasks across 7 categories (N=210), a diverse team with judge-based selection achieves a win rate of 0.810 against a single-model baseline, while a homogeneous team scores 0.512 -- near chance (Glass's Δ= 2.07). Judge-based selection outperforms MoA-style synthesis by Δ WR = +0.631 -- the synthesis approach is preferred over the baseline in zero of 42 tasks by the judge panel. A decoupled evaluation with independent judges confirms all directional findings (Spearman ρ= 0.90). Exploratory evidence suggests that including a weaker model improves performance while reducing cost (p < 10-4, not pre-registered). Our results suggest that selector quality may be a more impactful design lever than generator diversity in single-round generate-then-select pipelines.
Sources
- Mixture-of-Agents Enhances Large Language Model Capabilities
- Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial?
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
- Improving Factuality and Reasoning in Language Models through Multiagent Debate
- Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs
- When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent Reasoning
- LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion
- Optimizing Model Selection for Compound AI Systems
- Learning to summarize from human feedback
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- The Value of Variance: Mitigating Debate Collapse in Multi-Agent Systems via Uncertainty-Driven Policy Optimization
- More Agents Is All You Need
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Large Language Models are Inconsistent and Biased Evaluators
- LLM Evaluators Recognize and Favor Their Own Generations
- Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Related papers
- Highway Congestion Reduction through Reinforcement Learning Based Eulerian Headway Control
- You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents
- Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems
- MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimization
- PeroMAS: A Multi-agent System of Perovskite Material Discovery
- StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning