Mo' Models, Mo' Problems: How to best select model pools when designing Multi-Agent Systems
cs.MA, cs.AI
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: 8 pages main, 23 pages total. Accepted to REALM 2026 as part of EMNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multi-agent Systems (MAS) combine multiple model outputs to solve complex reasoning tasks.
Terminology
Abstract
Multi-agent Systems (MAS) combine multiple model outputs to solve complex reasoning tasks. However, despite rapid growth of available open-source models, there is limited research on how to select optimal model candidates out of this massive pool. We systematically evaluate 8 model selection strategies (including model size, accuracy and answer diversity) across before-generation (routing) and after-generation (majority-voting, LLM-as-a-judge) MAS architectures on challenging scientific benchmarks. Our findings show a significant gap between theoretical oracle potential and actual performance: Expanding candidate pool sizes often degrades performance below that of the top performing base-model. We find that candidate selection within a single model family is the strategy that yields the best relative performance over a standalone model. These results demonstrate that adding arbitrary models to a heterogeneous MAS can introduce system instability, highlighting model selection as a critical design choice for multi-agent systems.
Sources
- Intern-S1: A Scientific Multimodal Foundation Model
- Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process
- Harnessing Multiple Large Language Models: A Survey on LLM Ensemble
- cosmosage: A Natural-Language Assistant for Cosmologists
- The Llama 3 Herd of Models
- Tryage: Real-time, intelligent Routing of User Prompts to Large Language Models
- Towards a Science of Scaling Agent Systems
- LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing
- The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
- Routoo: Learning to Route to Large Language Models Effectively
- Olmo 3
- gpt-oss-120b & gpt-oss-20b Model Card
- Route-and-Reason: Scaling Large Language Model Reasoning with Reinforced Model Router
- Gemma 4 Technical Report
- Qwen3 Technical Report
- FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks
- Chem-R: Learning to Reason as a Chemist
- Bench-CoE: a Framework for Collaboration of Experts from Benchmark
- Understanding Agent Scaling in LLM-Based Multi-Agent Systems via Diversity
- ChemDFM-R: A Chemical Reasoning LLM Enhanced with Atomized Chemical Knowledge
Related papers
- Highway Congestion Reduction through Reinforcement Learning Based Eulerian Headway Control
- You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents
- Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems
- MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimization
- PeroMAS: A Multi-agent System of Perovskite Material Discovery
- StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning