From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing

arXiv:2606.06924 · cs.LG · Submitted 2026-06-05 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing".

Jane: The paper was written by Guannan Lai, Haoran Hu, Long Chen, Zhenguo Li and Han-Jia Ye from School of Artificial Intelligence, Nanjing University and National Key Laboratory for Novel Software Technology, Nanjing University and Hong Kong University of Science and Technology and Frontier Robotics.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Tom: So, we just wrapped up talking about moving from single outcomes to capability distributions using "From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing." Now, Jane, can you walk us through what the paper summarizes about the core mechanism they propose?

Jane: In simple terms, the paper’s summary points out that current methods are often too narrow; they only supervise on what's easy to test or what was explicitly provided in a small set of examples.

Lu: What that implies is that we're missing out on the vast middle ground—the areas where the model *should* perform well but hasn't been shown a specific example of doing so.

Meng: From an engineering standpoint, relying only on sampled outcomes means our performance metrics are inherently biased toward the data we *happened* to collect for testing, which is risky.

Lalam: The summary really highlights that this approach allows us to build AI systems that feel less brittle and more adaptable when they encounter novel situations in real life.

Jane: Exactly, it’s about capturing the *potential* rather than just confirming the *past* performance.

Tom: So, if I understand correctly, their summary shows that existing supervision techniques treat capability like a series of disconnected points on a graph?

Lu: Precisely; they're treating it discretely when nature and intelligence function continuously across many related tasks.

Meng: If we can model the distribution, we can create routing logic that says, "Given this input falls into this cluster of difficulty, use Model A because its distribution peaks here."

Jane: That sounds like a much smarter way to route queries than just sending everything to one massive black box model.

Lalam: This capability understanding fosters a kind of systemic intelligence; the AI doesn't just give an answer, it suggests *which* type of thinking is appropriate for the question asked.

Tom: It’s really about building self-aware routing, isn't it? Moving beyond simple keyword matching to deep functional mapping.

Lu: They are essentially proposing a way to supervise the *relationship* between different tasks, which is a level of abstraction most supervision methods ignore.

Meng: I wonder how computationally expensive it is to actually estimate these distributions across an entire model's capacity? That’s my practical hurdle right now.

Jane: We need to keep that computational cost in mind as we transition into discussing the actual improvements they suggest, because that’s where the engineering really gets interesting.

Improvements Suggested: Tom: Jane, we've talked about the gap—the difference between what current methods sample and what capability distributions suggest. Now, can you elaborate on the specific improvements "From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing" suggests?

Jane: The core improvement seems to be moving supervision targets from single outputs toward modeling the entire underlying probability space of possible behaviors for a given task family.

Lu: What I find exciting here is that they aren't just suggesting *calculating* the distribution; they are suggesting novel ways to *train* the model specifically to adhere to those probabilistic boundaries.

Meng: So, instead of just optimizing the loss function based on the ground truth answer, we’re optimizing it based on keeping our prediction within a constrained, theoretically sound distributional envelope?

Lalam: That implies a shift in what we value in AI: not just correctness on known tasks, but *predictable competence* across unknown task variations.

Tom: It sounds like they've given us the tools to supervise the *process* rather than just supervising the final product, which is a huge philosophical shift for LLMs.

Jane: Right, Tom

Paper discussion segment 3: Tom: So we’ve established that the biggest problem is that current methods treat model ability like a collection of isolated, noisy data points, but Jane, what does the paper suggest as a concrete improvement over this approach?

Jane: The authors propose using DARS, which stands for Distribution-Aware Routing Supervision. It lets you build supervision signals not just from one single response per query-model pair, but from an entire set of observed behaviors.

Lu: That’s where it gets fascinating because the theory shifts completely; we're not just looking at the average performance anymore, we’re mapping the full spectrum of what is possible for a task.

Meng: From an engineering standpoint, this means we can design routers that aren't just brittle. Instead of failing when they hit a query they haven’ve seen once, they know where it sits within a known cluster of difficulty and then route accordingly.

Lalam: That shift in capability understanding allows us to build systems that don't just solve problems, but understand the complexity inherent in solving them—it’s about recognizing the necessary mode of thought for a more profound cultural impact.

Tom: It sounds like it's not just about *how* we route, but *why* we are routing. It’s about giving the intelligence a better map of its own abilities.

Jane: Exactly, so she adds that it provides reliable labels because instability due to randomness is minimized by treating the samples as part of a larger distribution.

Lu: And I think this opens up such exciting possibilities for how we can design truly adaptive systems, where the routing logic itself learns from the probabilistic characteristics of tasks.

Meng: It definitely helps with deployment reliability too, by allowing us to predict not just performance but also model variability across different cost constraints.

Lalam: The ability to predict that models have inherent uncertainties allows us to allocate human and computational resources in a way that feels more intelligent and less waste-prone.

Tom: It's a huge leap from simply having "a good model" to having a much more nuanced understanding of the model’s entire performance profile.

Jane: And because Meng pointed out the cost aspect, it’s also about ensuring we aren't overpaying for a single powerful model when a slightly cheaper one is perfectly capable in most reliable scenarios.

Lu: Imagine that applying to other domains; instead of just classifying data, we are classifying the *potential* of the data itself.

Meng: Exactly, so he says it helps us choose the right tool for the job, not just because it works, but because we know exactly how robustly it works across its own observed distribution.

Lalam: This capability-centric routing leads to a future where AI is less of a black box and more of an intelligent partner that guides us toward better decisions.

Tom: It's clear the shift from single sample to understanding the entire distribution changes everything, but what do we need to consider next?

Conclusion: Tom: So, Jane, wrapping this up—it really feels like we’ve seen how much better things get when you move past just getting one sample answer and start mapping out whole capability ranges.

Jane: That’s exactly it, Tom; understanding that distribution rather than just optimizing for the average outcome is such a huge leap forward for building robust AI systems.

Lu: What strikes me most after hearing everyone talk is how this fundamentally changes our idea of "success" in LLMs; we aren't just aiming for the right answer, but mapping out *why* it might be right across many scenarios.

Meng: From a build standpoint, Lu makes sense that it shifts focus, but I keep wondering about the overhead—how much compute does mapping these distributions actually add compared to just training on curated hard examples?

Lalam: Meng touches on practicality, but I think the long-term cultural gain from this increased reliability outweighs any initial computational cost; giving people predictable AI assistance boosts trust across entire sectors.

Jane: It sounds like the consensus is that this approach to supervision tackles a core weakness in current models, moving us toward something much more dependable for real-world integration.

Tom: Absolutely, Jane; it’s about moving from anecdote to architecture, giving us a much clearer picture of what these models are actually capable of doing when pushed beyond the easiest prompts.

Lu: It really opens up possibilities for reasoning tasks where ambiguity is common, like complex scientific problem-solving that needs multiple pathways checked.

Meng: If we can better characterize those failure modes using capability distributions, then we can engineer guardrails that aren't just simple rule sets, but actual predictive safety nets.

Lalam: And the ability to communicate that uncertainty—to show the user *why* an answer might be less certain—that’s a huge step toward making AI feel genuinely helpful rather than just authoritative.

Tom: Right, so as we wrap up our discussion on "From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing," it seems this methodology is going to redefine how we supervise and deploy these powerful systems.

Jane: It’s been an incredible deep dive; thanks so much to all of you for walking us through the nuances of this research today. We'll catch up next time when we tackle another fascinating paper on the cutting edge!

School of Artificial Intelligence, Nanjing University · National Key Laboratory for Novel Software Technology, Nanjing University · Hong Kong University of Science and Technology · Frontier Robotics

cs.LG

Submitted: 2026-06-05

Updated: 2026-09-04

Comments: Accepted to EMNLP 2026 Main Conference

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 85/100

The gist: From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing The paper identifies a fundamental limitation in existing Large Language Model (LLM) routing methods: the

Key concepts

Capability Distributions
Instead of viewing a model's ability as isolated points, this concept models the entire range of possible behaviors for a task family. It captures the potential performance rather than just confirming past results.
Distribution-Aware Routing Supervision (DARS)
This proposed method allows supervision signals to be built from an entire set of observed behaviors, not just one response per query-model pair. It helps build routers that are less brittle and more adaptable.
LLM Routing
The process of directing a query to the most appropriate AI model or module. The paper suggests improving this by understanding the input's difficulty cluster and selecting a model whose capability distribution peaks there.

Terminology

Summary

From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing

The paper identifies a fundamental limitation in existing Large Language Model (LLM) routing methods: the reliance on single-shot supervision. Traditional routing protocols treat a model’s single response to a query as its capability label, which the paper argues is merely a noisy observation of a query-model pair’s behavior rather than a reliable capability estimate. This assumption introduces systematic noise into routing supervision, making learned routing policies less reliable.

Motivation and Problem Formulation

LLM generation is inherently stochastic. Even for the same original query, different sampled outcomes may suggest different model preferences. This variability in a model’s sampled response can alter its observed score; changes in these scores can change which model appears preferable for a query; and routers trained on such sample-dependent labels may learn policies that reflect incidental generation noise rather than stable differences in model capability.

The traditional routing objective is formulated as a constrained optimization problem:

sum i=1 N q i,r(x) s.t. sum i=1 N c i,r(x) C

where q i,m is the observed performance and c single is the observed cost of a single generated response. The paper notes that this implicitly treats one sampled generation as a reliable estimate of both behavior and cost.

The Proposed Solution: DARS

To address this issue, the authors propose DARS (Distribution-Aware Routing Supervision), a framework that constructs routing supervision from a distributional view of model behavior rather than relying on isolated sampled outputs. DARS captures uncertainty from two sides:

  1. Input-side uncertainty: This is captured by using semantically preserving prompt rewrites to capture sensitivity to query formulation, as LLMs can be sensitive to prompt formulations even if the underlying task remains the same.

  2. Output-side uncertainty: This is captured through repeated decoding (stochastic generation), acknowledging that a single generated response may not fully characterize how well a model handles a query.

DARS Workflow and Mechanism

DARS operates in four main stages:

  1. Repeated Observation Construction: For each training query x, DARS generates five semantically equivalent rewrites using GPT-4o. For every rewritten variant, it performs five independent stochastic decoding runs on each candidate model m. This results in a 5 times 5 observation matrix for every query-model pair (n, j), where each entry records both the response quality and the corresponding inference cost.

  2. Distributional Capability Estimation: DARS summarizes this observation set into three capability signals:

  • Expected quality: mu q(x, m) = Mean q m(n,j)

  • Expected cost: mu c(x, m) = Mean c(n,j) m

  • Performance risk (variability): sigma q(x, m) = Std q m(n,j)

  1. Risk-Aware Supervision Construction: DARS constructs a risk-aware utility:

U(x, m) = mu q(x, m) - lambda mu c(x, m) - beta sigma q(x, m

This utility incorporates expected performance (mu q), penalizes inference cost (lambda mu c), and penalizes unstable behavior (beta sigma q). The preferred model for query x is then determined by m*(x) = m in M U(x, m).

  1. Routing Policy Learning: This distribution-aware utility provides a unified supervision source for various routing methods (regression-based, utility-based, classification-based, etc.), decoupling supervision construction from router design.

Experimental Findings

The paper conducts a diagnostic analysis across three datasets—GPQA (scientific reasoning), MATH-500 (mathematical problem solving), and DROP-800 (reading comprehension)—to examine the limitations of single-shot supervision:

  • Observation 1: Single-shot labels are unstable. The results show that instability is not limited to response scores but propagates to model selection. For example, GPQA exhibits an outcome instability of 0.715 and a winner flip rate of 0.970.

  • Observation 2: Different tasks exhibit different uncertainty profiles. GPQA and MATH-500 are dominated by output-side uncertainty, while DROP-800 shows comparable input-side and output-side uncertainty, indicating that single-shot failure modes vary across tasks.

  • Observation 3: Single-shot supervision induces unstable learned routers. When training a router on 100 randomly sampled single-shot training sets, the resulting test performance distribution shows non-negligible policy variance, confirming that stochastic supervision propagates into the learned routing policy.

Performance and Conclusion

The main experiments compare DARS against various baseline routers (MLP, MIRT, EmbedLLM, etc.). The results show that DARS consistently improves almost all routers across the three datasets. This benefit is not tied to a specific router design but to the reliability of the supervision signal.

Furthermore, an ablation study confirms that both input-only and output-only variants improve upon single-shot baselines, but neither matches full DARS, confirming that input-side and output-side observations are complementary. The paper concludes that DARS provides more reliable routing signals and generally improves routing performance across diverse router families under cost-quality-risk trade-offs.

Improvements for AI systems

The core flaw in current LLM routing systems is the assumption that a single sampled response represents a reliable capacity estimate. We must replace this unreliable point-estimate supervision with Distribution-Aware Routing Supervision (DARS), which treats model behavior as an underlying capability distribution.

The following specific modifications should be implemented into the training pipeline of any LLM routing system:

1. Advanced Data Collection Protocol (Input and Output Uncertainty Capture):

Instead of sampling a single response y i,m for query x i and model m, we must construct a multi-view observation set O x,m.

  • Input-Side Perturbation: For every training query x, generate five semantically equivalent rewrites, x(1), x(2),, x(5), using a strong LLM (e.g., GPT-4o), ensuring these rewrites maintain the original task semantics and gold answers.

  • Output-Side Stochasticity: For each of the five rewritten queries, execute five independent stochastic decoding runs (M=5) under the specified sampling parameters (Temperature T=0.7, Top-P =0.95).

  • Result: This generates a 5 times 5 matrix of observations for every query-model pair, O x,m.

2. Statistical Aggregation and Feature Engineering: For each resulting observation set O x,m, we compute three critical distributional capability signals:

  • Expected Quality (mu q): Calculate the mean of all task-specific quality scores across all rewrites times decodings.

mu q(x, m) = Mean q n,j over all observations in O x,m

  • Expected Cost (mu c): Calculate the mean of all normalized inference costs.

mu c(x, m) = Mean c n,j over all observations in O x,m

  • Performance Risk (sigma q): Calculate the standard deviation of the quality scores to quantify performance variability.

sigma q(x, m) = Std q n,j over all observations in O x,m

3. Training Objective Transformation (Risk-Aware Utility):

Replace the traditional point-estimate utility U single = q(x, m) - lambda c(x, m) with the risk-aware utility U DARS. This objective incorporates performance variability (sigma q) as a penalty.

U(x, m) = mu q(x, m) - lambda mu c(x, m) - beta sigma q(x, m)

where lambda (cost coefficient) and beta (risk coefficient) are tuned hyperparameters. The the router is trained to maximize this utility: r*(x) = U(x, m).


By implementing DARS, the improved LLM routing system will achieve:

  1. Increased Routing Reliability: The model selection policy will be significantly more stable. It eliminates winner flip errors—where two models are equally good in a single sample but consistently ranked differently across multiple samples—by grounding the decision in the aggregate distribution of capability.

  2. Robust Performance Gains: The system will consistently outperform routers trained on single-shot data, particularly on datasets (like GPQA) where single-shot label instability is high, leading to a more accurate mapping between query difficulty and optimal model selection.

  3. Enhanced Risk Management: The explicit inclusion of sigma q allows the router to penalize models that are highly erratic or unpredictable for a given task, even if their average performance (mu q) is high. This leads to safer, more consistent deployment decisions than simple average-based routing.

  4. Optimized Cost-Quality Trade-Off: The system will demonstrate a superior Pareto frontier in cost-quality trade-off analysis (as shown in Figure 4b), identifying scenarios where a slightly less capable but highly stable and cost-effective model is preferable to an erratic, expensive alternative.

Abstract

Existing LLM routing methods often construct supervision from a single sampled response for each query--model pair. Because LLM generation is stochastic, however, such an observation can be an unstable estimate of model capability: semantically equivalent query formulations and repeated decoding may yield different scores and even different model preferences. We show that this instability can further propagate from routing labels to learned routing policies. To address this issue, we propose DARS (Distribution-Aware Routing Supervision), which estimates query-level model capability from repeated observations spanning semantics-preserving query rewrites and stochastic decoding. DARS summarizes expected quality, expected cost, and performance variability to construct risk-aware supervision without changing the downstream router architecture. Experiments across diverse tasks and routing methods show that DARS generally improves routing utility and cost--quality trade-offs over single-shot supervision. Further analyses show that its benefits persist under moderate sampling budgets and different decoding temperatures. These results suggest that reliable LLM routing should move beyond individual sampled outcomes and instead model query-level capability distributions.

Sources

Related papers