What Does a Routing Oracle Measure Under Stochastic Decoding? Coupling, Scorer Choice, and Single-Commit Ceilings
cs.LG
Submitted: 2026-07-03
Updated: 2026-09-27
Comments: Substantially revised and retitled; supersedes empirical summaries and interpretations in earlier versions. Coupling-aware fixed-archive audit, retrospective policy illustration, and limited reference-based human check. 19 pages, 8 figures. Source and evidence: https://github.com/luka-krixvon/routing-oracle-experiment/tree/main/releases/2026-09-27
Code: https://github.com/luka-krixvon/routing-oracle-experiment
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- A Unified Approach to Routing and Cascading for LLMs
- Dynamic Model Routing and Cascading for Efficient LLM Inference: A Survey
- RouterBench: A Benchmark for Multi-LLM Routing System
- LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing
- Route-and-Reason: Scaling Large Language Model Reasoning with Reinforced Model Router
- Best-of-$\infty$ -- Asymptotic Performance of Test-Time LLM Ensembling
- Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
- The Capability Frontier: Benchmarks Miss 82% of Model Performance
- Unsolvability Ceiling in Multi-LLM Routing: An Empirical Study of Evaluation Artifacts
- Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial?
- RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs
- Within-Model vs Between-Prompt Variability in Large Language Models for Creative Tasks
- Measuring all the noises of LLM Evals
- On Randomness in Agentic Evals
- Quantifying Variance in Evaluation Benchmarks
- When Routing Collapses: On the Degenerate Convergence of LLM Routers
- Expected Reward Prediction, with Applications to Model Routing
- LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks