Dynamic Model Routing and Cascading for Efficient LLM Inference: A Survey
cs.NI, cs.CL, cs.PF
Submitted: 2026-02-23
Updated: 2026-08-30
Comments: Accepted by TMLR (2026). Work funded by ADAPT Centre, Trinity College Dublin, and Huawei Ireland
Journal ref: Transactions on Machine Learning Research (TMLR). 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: The rapid growth of large language models (LLMs) with diverse capabilities, costs, and domains has created a critical need for intelligent model selection at inference time.
Terminology
Abstract
The rapid growth of large language models (LLMs) with diverse capabilities, costs, and domains has created a critical need for intelligent model selection at inference time. While smaller models suffice for routine queries, complex tasks demand more capable models. However, static model deployment does not account for the complexity and domain of incoming queries, leading to suboptimal performance and increased costs. Dynamic routing systems that adaptively select models based on query characteristics have emerged as a solution to this challenge. This survey provides a systematic analysis of multi-LLM routing and cascading approaches, focusing on systems that route queries across a pool of independently trained LLMs at inference time. We cover diverse routing paradigms, including query difficulty, human preferences, clustering, uncertainty quantification, reinforcement learning, multimodality, and cascading. For each paradigm, we analyze representative methods and examine key trade-offs. Beyond taxonomy, we introduce a conceptual framework that characterizes routing systems along three dimensions: when decisions are made, what information is used, and how they are computed. This perspective highlights that practical systems are often compositional, integrating multiple paradigms under operational constraints. Our analysis demonstrates that effective multi-LLM routing requires balancing competing objectives. Choosing the optimal routing strategy depends on deployment and computational constraints. Well-designed routing systems can outperform even the most powerful individual models by strategically leveraging specialized capabilities across models while maximizing efficiency gains. Meanwhile, open challenges remain in developing and evaluating routing mechanisms that generalize across diverse architectures, modalities, and applications.
Sources
- Qwen Technical Report
- Models Under SCOPE: Scalable and Controllable Routing via Pre-hoc Reasoning
- Evaluating Large Language Models Trained on Code
- LLM Routing with Dueling Feedback
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- Training Verifiers to Solve Math Word Problems
- The Llama 3 Herd of Models
- MMR-Bench: A Comprehensive Benchmark for Multimodal LLM Routing
- DeBERTa: Decoding-enhanced BERT with Disentangled Attention
- LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing
- WebGPT: Browser-assisted question-answering with human feedback
- MetaLLM: A High-performant and Cost-efficient Dynamic Framework for Wrapping LLMs
- Proximal Policy Optimization Algorithms
- Route-and-Reason: Scaling Large Language Model Reasoning with Reinforced Model Router
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- CP-Router: An Uncertainty-Aware Router Between LLM and LRM
- Arch-Router: Aligning LLM Routing with Human Preferences
- ReLope: KL-Regularized LoRA Probes for Multimodal LLM Routing
- GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
Related papers
- HiFiNet: Hierarchical Fault Identification in Wireless Sensor Networks via Edge-Based Classification and Graph Aggregation
- Embodied AI in 6G Networks: From Intelligent Connectivity to Physical Intelligence
- Lightweight GenAI for Network Traffic Generation: Fidelity, Augmentation, and Classification
- EdgePoW: Adaptive Ingress-Aware Defense with Non-Interactive PoW Against Volumetric SYN Floods
- SoK: Where Do Flow Labels Come From? Auditing Label Provenance in Encrypted Traffic Benchmarks
- What is Normal? A Big Data Observational Science Model of Anonymized Internet Traffic