Mo' Models, Mo' Problems: How to best select model pools when designing Multi-Agent Systems

arXiv:2609.17306 · cs.MA, cs.AI · Submitted 2026-09-15 · Read on arXiv

cs.MA, cs.AI

Submitted: 2026-09-15

Updated: 2026-09-15

Comments: 8 pages main, 23 pages total. Accepted to REALM 2026 as part of EMNLP 2026

License: http://creativecommons.org/licenses/by/4.0/

The gist: Multi-agent Systems (MAS) combine multiple model outputs to solve complex reasoning tasks.

Terminology

Abstract

Multi-agent Systems (MAS) combine multiple model outputs to solve complex reasoning tasks. However, despite rapid growth of available open-source models, there is limited research on how to select optimal model candidates out of this massive pool. We systematically evaluate 8 model selection strategies (including model size, accuracy and answer diversity) across before-generation (routing) and after-generation (majority-voting, LLM-as-a-judge) MAS architectures on challenging scientific benchmarks. Our findings show a significant gap between theoretical oracle potential and actual performance: Expanding candidate pool sizes often degrades performance below that of the top performing base-model. We find that candidate selection within a single model family is the strategy that yields the best relative performance over a standalone model. These results demonstrate that adding arbitrary models to a heterogeneous MAS can introduce system instability, highlighting model selection as a critical design choice for multi-agent systems.

Sources

Related papers