The Routing Plateau: Understanding the Accuracy Limits of LLM Routers
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "The Routing Plateau".
Tom: Many current LLM routing methods converge to a narrow performance range far below the oracle router, indicating fundamental limits in their ability to handle query-specific routing decisions.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: Okay, moving on to the specifics of the paper, this is "The Routing Plateau: Understanding the Accuracy Limits of LLM Routers," and it’s authored by Yifan Lu, Qiyue Zhang, Shenrun Zhang, Zhibo Yu, Zhuang Wang, Hanjie Chen, and Jiarong Xing. These researchers are really laying out why we see this plateau in routing performance.
Jane: The title itself tells us the core issue: there's a ceiling to how good these AI routers can get because of a predictability bottleneck that they need to overcome. It’s not just about making the current methods slightly better, but fundamentally changing what the router learns to predict.
Lu: The authors set up this study by evaluating twenty-one different routing methods across five distinct benchmarks, and their main finding is that these diverse approaches all end up converging at a very similar accuracy level that stays significantly below the oracle router.
Meng: So, they found that even with twenty-one different designs—from clustering to learned classifiers—they aren't achieving the high accuracy we’d expect for instance-specific tasks. What does this mean for us when we look at developing new routing architectures?
Lalam: It suggests that simply tweaking the architecture of a router isn't enough; you need to address the core signal problem they identified, which is how these systems are learning to make decisions.
Tom: Exactly, and their analysis focuses on three specific observations they found: similar top-end accuracy, strong kNN-style routers that stay competitive, and a persistent gap between the best router and the oracle router. It really paints a picture of stagnation in this area of AI research.
The paper's summary: Jane: So, to summarize what the authors uncovered in "The Routing Plateau: Understanding the Accuracy Limits of LLM Routers," they identified that current routers struggle because they learn coarse, global patterns about model performance instead of the fine-grained signals needed for specific queries.
Lu: That's the crux of it; they found that these routers mainly capture averaged trends in model capability rather than the subtle differences in correctness for a particular input. The set of models that are correct for many queries doesn't always align with overall model performance, which confuses the learning process.
Meng: From an engineering viewpoint, if they learn global patterns, it implies that if we feed them enough data of average performance, they will only ever be good at average cases and fail when things get specific. How does that translate into building a system that handles rare but important queries?
Lalam: It means the current AI systems are excellent at handling the majority of common requests but brittle when faced with the unique, hard instances that really stress model selection. This makes our service less robust overall.
Tom: And this is where it gets critical for query difficulty; they stratified queries into hard and easy subsets based on whether the set of correct models varies a lot, showing that these hard queries make up sixty-nine point nine percent to ninety point eight percent of the gap between routers and the oracle router.
Jane: That high percentage confirms that the shared failure point is exactly where it matters most—when we need instance-specific routing decisions rather than just relying on a general model ranking. It really shows that coarse estimates are not cutting it for those tough scenarios.
The paper's improvements: Lu: To break this plateau, the authors propose three main levers: increasing training datasets, building stronger query encoders, and implementing end-to-end fine-tuning of the router itself. They show that combining these elements leads to a combined accuracy gain of up to two point one three percentage points for today's routers.
Meng: That two point one three percentage point gain sounds substantial; it suggests that focusing on data, better representations, and training the whole system together actually yields noticeable results in closing that oracle gap. What does that imply for our development roadmap?
Lalam: It means the path forward isn't just about making one component smarter; it’s about a coordinated effort across multiple stages of development to improve routing performance holistically.
Tom: The ablation study shows that while individual changes like scaling data or using a larger encoder give modest gains, the real improvement comes when you combine data scaling with fine-tuning and using the larger encoder on top, which gives a total gain of one point two four percentage points over the baseline.
Jane: So, we're looking at a strategy where we need more examples to teach it what's right, better input features to understand the query better, and then training the whole router end-to-end on that data. It sounds like a multi-pronged approach is necessary here.
Conclusion: Tom: So, wrapping up this discussion on "The Routing Plateau: Understanding the Accuracy Limits of LLM Routers," the paper clearly shows that we're hitting a limit because current routers rely too heavily on global trends instead of instance-specific signals. The authors show that scaling training data, using larger encoders, and end-to-end fine-tuning can move us up to two point one three percentage points in accuracy.
Jane: While those gains are encouraging for improving the performance of these routing systems today, the authors are upfront that this approach doesn't quite close the gap entirely because solving a query-only prediction problem seems fundamentally difficult at this level.
Lu: They suggest that future progress will require moving beyond static query embeddings and incorporating richer instance-specific evidence, like model-pool-aware objectives or lookahead signals from partial generation to get better hints.
Meng: From an engineering side, I think the immediate practical safeguard is focusing on per-subgroup evaluation at deployment, as scaling traffic can concentrate usage on a few winners and leave specialized models behind in niche areas.
Lalam: I agree with that operational safety measure; it’s crucial to ensure that our deployment strategies account for the limitations they identified, keeping things fair across different query types.
Tom: So, in short, "The Routing Plateau: Understanding the Accuracy Limits of LLM Routers" gives us a clear map on where we are stuck and what specific steps—scaling data, upgrading encoders, and fine-tuning end-to-end—can take us further toward better model selection. That’s all for this episode.
Rice University
cs.LG
Submitted: 2026-05-27
Updated: 2026-09-28
Comments: 29 Pages, 23 Tables, 10 Figures
Code: https://github.com/Not-Diamond/RoRF
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 83/100
The gist: Many current LLM routing methods converge to a narrow performance range far below the oracle router, indicating fundamental limits in their ability to handle query-specific routing decisions.
Key concepts
- Routing Plateau
- This phenomenon describes how many different router designs achieve nearly identical accuracy ceilings, staying significantly below the best possible performance (the oracle router). It shows that current methods are fundamentally limited in their ability to make precise routing decisions based on individual query details.
- Correctness-Prediction Bottleneck
- Routers struggle because they must learn to guess which model will provide the correct answer for a specific question. Current methods only learn general, coarse patterns about model capability instead of fine-grained signals that distinguish correct answers for unique instances.
- Hard Queries
- These are difficult queries where the set of models capable of answering correctly varies significantly, often requiring the selection of a single correct model. These hard queries account for a large portion of performance gaps because routers lack the instance-specific knowledge needed to handle them reliably.
Terminology
Summary
Many current LLM routing methods converge to a narrow performance range far below the oracle router, indicating fundamental limits in their ability to handle query-specific routing decisions. This study evaluates 21 diverse routing methods across five benchmarks and reveals that this routing plateau
is caused by a correctness-prediction bottleneck where routers learn coarse, global patterns instead of fine-grained, instance-level signals.
The Routing Plateau Phenomenon
The central finding is the existence of a routing plateau,
where many different router designs, including kNN-style methods, achieve very similar accuracy ceilings that remain significantly below the oracle router. This phenomenon is characterized by three specific empirical observations: (1) Similar top-end accuracy,
where the top 15 routers differ by only a small margin; (2) Strong kNN-style routers,
which remain competitive across benchmarks; and (3) a persistent oracle gap,
as all existing routers trail the oracle by 10–30 percentage points.
The Correctness-Prediction Bottleneck
The plateau stems from the realization that routers must learn to infer which models are likely to answer a given query correctly, but current methods often fail to reliably learn this signal from available data and query representations. Specifically, current routers tend to learn coarse, global patterns of model capability rather than fine-grained, instance-level differences in correctness.
This occurs because the set of models that produce correct answers for many queries varies in ways not always aligned with overall model capability; higher-average-performing models are not consistently the most reliable on every instance. Consequently, their shared reliance on global-model-capability signals causes them to converge to a narrow accuracy ceiling far below that of the oracle router.
Analysis of Query Difficulty
The study stratifies solvable queries into hard and easy subsets based on whether the set of correct models varies significantly. A query is called hard
if it requires selecting the only correct model, or if there are multiple potentially correct models but the global top-1 model is incorrect on that instance. The analysis shows that hard queries make up only 11.1%–35.4% of each benchmark, yet account for 69.9%–90.8% of the overall gap.
This disparity confirms that routers share the failure of routing hard queries because they rely on coarse estimates rather than instance-specific correctness prediction needed for these difficult cases.
Directions for Breaking the Plateau
To move beyond this limit, the paper explores three key levers: (1) larger training datasets,
(2) stronger query encoders,
and (3) end-to-end fine-tuning.
The authors constructed a new large-scale routing training dataset with 300k queries and 2.8 million query–model correctness labels. By scaling the training set, upgrading the query encoder from MODERNBERT-BASE to MODERNBERT-LARGE, and fine-tuning the router end-to-end, they obtained a combined accuracy gain of up to 2.13 pp,
which closes approximately 14.6% of the oracle gap.
Synergy of Scaling and Fine-Tuning
The ablation study demonstrates that individual changes provide only modest gains, but the full combination yields the largest improvement. The dominant gain appears when data scaling is combined with fine-tuning: moving from 30k FT with MODERNBERT-BASE to 300k FT with the same encoder improves accuracy by 0.77 pp.
Furthermore, Adding the larger encoder on top gives the best result, Scaled-FT,
yielding a total gain of 1.24 pp over the baseline. This suggests that data scaling, encoder scaling, and task-specific fine-tuning are complementary levers for moving beyond the current frozen-encoder plateau.
Conclusion and Future Work
The proposed Scaled-FT narrows the gap to the oracle but does not close it, as this gap is expected due to the fundamental difficulty of solving a query-only prediction problem. The authors suggest that future progress requires "richer instance-specific evidence beyond static query embeddings, such as model-pool-aware objectives that compare models jointly (e.g., pairwise loss function), lightweight lookahead signals from partial generation or confidence estimates, and richer query representations. They conclude that the most actionable near-term safeguards involve
per-subgroup evaluation, transparent disclosure of pool composition at deployment, and open release of routing benchmarks."
The gist: many current LLM routing methods converge to a narrow performance range far below the oracle router. This phenomenon is caused by a correctness-prediction bottleneck where routers learn coarse, global patterns instead of fine-grained, instance-level signals. The authors show that scaling training data, strengthening query encoders, and end-to-end fine-tuning can improve accuracy by up to 2.13 pp for today’s routers.
Improvements for AI systems
Here are specific, actionable improvements for AI systems derived from the findings of this paper, categorized by the bottleneck they address:
)1. Move Beyond Coarse Global Patterns to Fine-Grained Instance-Specific Routing (Addressing the Correctness Prediction Bottleneck):
The core finding is that current routers learn coarse, global patterns instead of fine-grained, instance-level correctness signals.
Current routers tend to learn coarse, global patterns of model capability rather than fine-grained routing signals. For many queries, however, the set of models that produce correct answers varies in ways that are not always aligned with overall model capability.
Specific improvements:
Implement a novel routing objective that explicitly penalizes reliance on top-performing models when the query is classified as
hard(i.e., instance-specific routing required). This could involve training the router to maximize accuracy on hard queries while simultaneously minimizing the divergence between its selection distribution and a ground truth distribution that favors models correct for that specific query instance, rather than just maximizing average performance.
Develop a mechanism to inject explicit
instance-specific correctness prediction signalsinto the router's loss function. This requires moving beyond simple cross-entropy loss on binary labels (Ym) to incorporate features or representations that capture the subtle differences in model behavior for specific query embeddings, effectively forcing the model to learnwhich model is correct for this specific input,not justwhich model is generally best.
Introduce a
model-pool-aware objectiveduring training. This objective should compare models jointly (e.g., using pairwise loss functions) rather than treating them as independent entities whose performance is aggregated into a single global score. This would help the router learn to exploit nuanced trade-offs between models that are only correct on specific subsets of queries, which is critical for hard queries where thehighest average-performing modelfails.
- Enhance Query Representation Power (Addressing the Encoder Limitation):
The paper suggests that stronger query encoders can provide richer signals to overcome the plateau.
We further study how to move beyond the plateau and find that larger training datasets, stronger encoders, and end-to-end fine-tuning can further improve routing accuracy.
Specific improvements:
Systematically upgrade the query encoder from current baselines (e.g., MODERNBERT-BASE) to larger models (e.g., MODERNBERT-LARGE or similar high-capacity representations). This richer representation should capture more complex semantic and structural features of the query, providing the router with more discriminative information to distinguish between models that might otherwise appear similar based on coarse embeddings.
Explore multi-encoder fusion techniques where a single query is processed by multiple encoders (e.g., sentence-transformer and BERT). The resulting fused representation could be used as input to the router, capturing complementary signals that enhance the ability to identify fine-grained routing differences across models.
- Leverage Data Scaling and End-to-End Fine-Tuning (Addressing Training Limitations):
The most significant gains (up to 2.13 pp) were achieved by scaling data, using larger encoders, and end-to-end fine-tuning.
By scaling the routing training set from 30k to 300k queries, upgrading the query encoder from MODERNBERT-BASE (∼110M parameters) to MODERNBERT-LARGE (∼340M parameters), and fine-tuning the router end-to-end, we obtain a combined accuracy gain of up to 2.13 pp.
Specific improvements:
Implement a robust data augmentation and scaling pipeline for routing datasets, targeting 300k+ query instances with high-quality model correctness labels. This large dataset is essential for providing the dense supervision needed to learn fine-grained signals and should be used as the primary training source for any new router architecture.
Adopt end-to-end fine-tuning (FT) for all novel routing methods, moving away from fixed prediction heads on frozen embeddings. End-to-end FT allows the router to adapt its selection strategy directly to the downstream task performance and query representations, maximizing the utility of the scaled data and stronger encoders.
- Operational Deployment Safeguards (Addressing Real-World Risks):
The discussion in Section 6.4 points out risks associated with deployment:
At deployment, scaling this concentrates traffic on a few “winning” providers, marginalizes specialist models that excel on minority query distributions—non-English queries, niche domains, atypical formatting—exactly the hard regions where current routers already fail.
Specific improvements:
Integrate
model-pool-aware objectivesdirectly into the deployment and monitoring phase. This involves continuously auditing the router's performance not just on aggregate accuracy, but on metrics that track performance across different model subgroups (e.g., language, domain specialization). If a specific subgroup of queries consistently performs poorly under the current routing strategy, an adaptive mechanism should trigger a re-calibration or adjustment to the routing weights for that subgroup.
Develop lightweight generation-time signals (lookahead) to supplement static query embeddings. Instead of relying solely on pre-computed features, integrate minimal confidence estimates or partial generation feedback during the inference phase to allow the router to make more dynamic, real-time corrections when it detects an instance where its current coarse prediction is likely erroneous.
Abstract
LLM routing has become a popular approach to improve the cost-quality trade-off of LLM services by adaptively selecting a model for each query. Recent work has explored a broad range of routing methods, including clustering-based routers, learned classifiers, pairwise ranking, and confidence-based approaches. Our extensive study of 21 routing methods across five benchmarks reveals a consistent phenomenon that we call the routing plateau (Fig. 1): many methods, including kNN, achieve very similar accuracy and converge to a narrow performance range that remains far below the oracle router. Our analysis supports a correctness-prediction bottleneck hypothesis: current routers primarily learn global-average model performance trends rather than fine-grained, query-specific routing signals. As a result, they collectively fail on queries that require instance-specific routing decisions. Moreover, to understand whether the plateau can be alleviated with a better training setup, we construct a 300K-query benchmark (Nine-by-300k). More data, larger encoders, and end-to-end fine-tuning improve eight routers by 1.24 pp on average, but leave the plateau largely intact. These findings suggest that further progress may require inputs beyond the query itself, such as partial output trajectories that reveal how models attempt the task.
Sources
- Program Synthesis with Large Language Models
- MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
- No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
- Evaluating Large Language Models Trained on Code
- RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language Models
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
- SearchQA: A New Q&A Dataset Augmented with Context from a Search Engine
- GraphRouter: A Graph-based Router for LLM Selections
- Prompt-to-Leaderboard
- RouterBench: A Benchmark for Multi-LLM Routing System
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- Universal Model Routing for Efficient LLM Inference
- When Routing Collapses: On the Degenerate Convergence of LLM Routers
- LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing
- Rethinking Predictive Modeling for LLM Routing: When Simple kNN Beats Complex Learned Routers
- OptLLM: Optimal Assignment of Queries to Large Language Models
- Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models
- RouteLLM: Learning to Route LLMs with Preference Data
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks