Error-Aware Reverse Auction Mechanism for Large Language Model Routing
Shenzhen International Center for Industrial and Applied Mathematics · Shenzhen Research Institute of Big Data · The Chinese University of Hong Kong, Shenzhen · Shenzhen Campus of Sun Yat-sen University · Shenzhen Loop Area Institute
cs.GT, cs.AI
Submitted: 2026-08-13
Updated: 2026-09-28
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
Terminology
Summary
arXiv: 2608.12719v1 [cs.GT] 13 Aug 2026
The paper addresses the challenge of routing each query to a cost-effective large language model (LLM) to balance quality and cost. The authors identify two structural limitations of existing centralized routing paradigms:
-
Information–risk mismatch:
the task center bears the risk of failure but has less information about the LLMs, whereas providers hold richer private knowledge.
-
Scalability bottleneck:
the center must profile or train for every new model, incurring per-model overhead even for training-free routers, which hinders rapid expansion.
The authors propose a market-based routing paradigm that transfers ex-ante prediction to LLM providers via a reverse auction
(Figure 1). In this paradigm:
-
The task center acts as the buyer and solicits bids from providers, who report their self-predicted success probabilities and costs
-
the center maintains only a model-agnostic ex-post evaluator
-
This distributed design aligns information with risk and removes the need for model-by-model profiling for the buyer, thereby improving scalability
The key theoretical contribution is explicitly modeling what the authors call the inherent Dual Error
:
-
Providers' ex-ante success estimates are subjective and noisy
-
The center's ex-post evaluation is imperfect
Formally, the paper defines:
-
Ex-post evaluation error:
The buyer evaluates the output via the error-involved ex-post acceptance probability hi = σ(ϕi + εpost), where the evaluation error εpost has E[εpost] = apost and Var(εpost) = bpost
-
Ex-ante prediction error:
seller i predicts the buyer's acceptance probability... We introduce an independent prediction error εante,i with E[εante,i] = aante,i and Var(εante,i) = bante,i, and define the aggregated error ηi = εpost + εante,i
The seller's ex-ante subjective probability is gi = σ(ϕi + ηi), where E[ηi] = aη,i = apost + aante,i and Var(ηi) = bη,i = bpost + bante,i.
The mechanism works as follows (Algorithm 1):
-
Bidding:
Providers submit ex-ante bids
with self-predicted success probabilities and costs -
Allocation:
the center allocates the query
based on reported surplus scores ŝi = V gi − ci -
Execution:
The selected model answers
-
Evaluation and Payment:
The buyer evaluates the output and settles payment
with payment rule rj = V µ̃j − H, where H is the runner-up score
The payment rule is designed so that payments depend on the evaluator signal µ̃j ∈ 0, 1, a noisy proxy for the unobserved ground truth µj
rather than self-reported probabilities, preventing sellers from inflating their reported p̂i.
The paper establishes several key theoretical properties:
Theorem 3.2: Seller i's interim expected utility Uiseller(ŝi) is maximized at ŝi = T̄i. Hence, reporting ŝi = T̄i is a Bayesian best response.
Theorem 3.3: At the equilibrium ŝ = T̄, every seller satisfies IR: E[Uiseller θ̃i] ≥ 0.
Theorem 3.4: Two sufficient conditions guarantee CR:
-
(A) E[H] ≥ V∆gate, where ∆gate = Lσ√(bpost + a2post)
-
(B) ∆cons = E[p(1) − h(1)] ≥ 0
Theorem 3.7: The expected welfare loss satisfies 0 ≤ E[Wi⋆] − E[Wi†] = E[(V pi⋆−ci⋆) − (V pi†−ci†)] ≤ 2V Lσ(Mpost + Mante)
where Mpost = √(bpost + a2post) and Mante = maxi √(bante,i + a2ante,i).
The paper reveals three important robustness effects:
Proposition 3.8: If (gi − hi)(hi − pi) ≤ 0, for i ∈ i⋆, i†, then the welfare-loss bound tightens to 2V Lσ max Mante, Mpost.
This means when ex-ante and ex-post errors have opposite signs, they partially cancel out, reducing the resulting welfare loss.
Proposition 3.9: For link functions with vanishing tails (e.g., logistic), the deviation ∆(ϕi) = E[σ(ϕi + ϵ)] − σ(ϕi) satisfies limϕi→∞ ∆(ϕi) = 0.
This means flat-tailed links such as the logistic function saturate when ϕi is large, so score noise barely affects σ(ϕi) and the mechanism is robust in clear-cut cases.
Proposition 3.10: Injecting additional independent noise weakly reduces the maximal sensitivity of both ϕ ↦ gi(ϕ) and ϕ ↦ hi(ϕ).
This limits the return to strategic manipulation.
The paper simulates N = 5 heterogeneous sellers with V = 20 and d = 1.5, using a highly competitive market with a small top-two surplus gap (∆s = 0.44), where mild noise can flip the ranking.
Results show:
-
Figure 5a:
EA-RAM closely tracks this bound and maintains a stable welfare gap as σpost increases, whereas Error-Naive deteriorates sharply
-
Figure 5b:
EA-RAM consistently outperforms Error-Naive, showing that internalizing the evaluation mechanism avoids severe misallocation and yields more graceful degradation
The paper evaluates on RouterBench, which aggregates per-query accuracy and monetary cost for 11 LLMs,
spanning open-source (Llama, Mixtral) and proprietary models (GPT, Claude), across six benchmarks: HellaSwag, Winogrande, ARC-Challenge, MBPP, GSM8k, and MMLU.
Baselines compared: EmbedLLM, IRT-Router, RouteLLM, FrugalGPT, and Cascade Routing.
Key results:
-
Table 2: EA-RAM achieves the best AIQ (Average Improvement in Quality) across all benchmarks, with the base setting outperforming all baselines
-
Figure 6:
EA-RAM already outperforms baselines in the base setting, and seller-side local information further shifts the frontier upward and leftward
-
Figure 7: In noisy LLM-as-a-Judge evaluation experiments,
EA-RAM consistently performs best across all α, incurring the smallest welfare loss
-
Figure 8:
The center-side latency of EA-RAM remains nearly constant as the model pool grows, since the center only performs evaluator and auction computation and does not require retraining or per-model prediction
-
Table 3: Communication overhead
grows roughly linearly with N, while computation latency stays nearly constant,
with communication becoming the bottleneck only below 0.338–0.384 MB/s per channel
The paper's main contributions are:
-
"We propose EA-RAM, a market-based routing framework that shifts ex-ante prediction to LLM providers, addresses centralized routers' information–risk mismatch and per-model profiling bottleneck, and models LLM routing as an auction under explicit Dual Error."
-
"We establish theoretical guarantees under Dual Error: EA-RAM satisfies BIC and IR, admits sufficient conditions for CR, and enjoys a welfare-loss bound; we further reveal robustness insights, including error compensation, saturation stability, and noise-induced flattening."
-
Extensive experiments on simulations and real-world benchmarks show that EA-RAM is robust to Dual Error and improves economic efficiency compared to state-of-the-art centralized baselines.
The paper acknowledges several limitations:
-
Our present theory does not cover payment-rule variants such as bounded penalties or non-negative payments
-
in communication-heavy settings such as image- or video-based tasks, or under degraded network conditions, communication may become a more significant bottleneck
-
the auction protocol requires broadcasting each query to all candidate providers before selection, which may expose user prompts to non-winning providers and raises security concerns
-
providers may be incentivized to optimize toward the evaluator rather than the user's true underlying need
Future directions include: communication-efficient reverse-auction mechanisms,
extending the theory to bounded-penalty or non-negative-payment variants,
tightening the current welfare-loss bound,
generalizing the framework beyond single-winner routing to support multi-provider cross-checking, cascading, and multi-turn settings,
extending the model to heterogeneous risk preferences,
and providing more principled support for non-binary evaluation tasks.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems:
-
Improvement: Replace centralized router prediction with a reverse auction where LLM providers bid with self-reported success probabilities and costs.
-
Capability: The AI system can dynamically select the most cost-effective model for each query without requiring the central system to profile or train on every new model. This enables seamless scaling as new LLMs are added to the pool.
-
Improvement: Explicitly model both ex-ante prediction errors (from providers) and ex-post evaluation errors (from the center) when making routing decisions.
-
Capability: The AI system can quantify uncertainty in both provider claims and its own evaluation, leading to more robust model selection that degrades gracefully under noisy conditions rather than failing catastrophically.
-
Improvement: Use payment rules based on noisy evaluator signals rather than self-reported probabilities, with payments tied to the runner-up score.
-
Capability: The AI system can incentivize truthful reporting from providers, preventing gaming of the routing system and ensuring that providers are rewarded based on actual performance rather than inflated claims.
-
Improvement: Detect when provider overestimation and center underestimation (or vice versa) partially cancel out, and adjust routing confidence accordingly.
-
Capability: The AI system can identify scenarios where errors naturally balance, allowing it to make higher-confidence routing decisions even when individual error sources are significant.
-
Improvement: Recognize that when model quality signals are extreme (very high or very low), noise has minimal impact on routing decisions.
-
Capability: The AI system can make rapid, confident routing decisions for queries where model suitability is obvious, while allocating more computational resources to ambiguous cases where noise could flip the ranking.
-
Improvement: Understand that adding controlled noise to evaluation signals reduces the sensitivity of outcomes to strategic manipulation.
-
Capability: The AI system can intentionally inject calibrated noise into its evaluation process to reduce the return on adversarial provider behavior, making the system more robust to gaming.
-
Improvement: Keep the central evaluation component independent of specific model architectures, focusing only on output quality.
-
Capability: The AI system can evaluate any new model's output without retraining, enabling instant integration of new LLMs into the routing pool with near-zero marginal cost.
-
Improvement: Balance the linear communication overhead of broadcasting queries to all providers against the constant computation latency of the center.
-
Capability: The AI system can operate efficiently in bandwidth-constrained environments by understanding when communication becomes the bottleneck (below 0.34 MB/s per channel) and adapting its routing strategy accordingly.
-
Improvement: Use the theoretical welfare-loss bound (≤ 2V Lσ(Mpost + Mante)) to guarantee that routing decisions remain economically efficient even with noisy predictions.
-
Capability: The AI system can provide formal guarantees on the maximum economic inefficiency introduced by prediction errors, enabling deployment in cost-sensitive enterprise applications where worst-case performance matters.
-
Improvement: Extend the single-winner auction framework to support cascading, cross-checking, and multi-turn interactions between multiple providers.
-
Capability: The AI system can orchestrate complex workflows where multiple models collaborate, verify each other's outputs, or iteratively refine responses, while maintaining the same incentive-compatibility guarantees.
Abstract
Routing each query to a cost-effective large language model (LLM) is critical for balancing quality and cost, yet most routers rely on a centralized task center to predict model performance, creating an information-risk mismatch and a scalability bottleneck as the model pool grows. We propose a market-based routing paradigm that shifts ex-ante prediction to LLM providers via a reverse auction, where providers bid with self-predicted success probabilities and execution costs. To account for inherently noisy provider predictions and center evaluations, we introduce the Error-Aware Reverse Auction Mechanism (EA-RAM), which explicitly models this inherent Dual Error. We prove that EA-RAM is Bayesian incentive compatible and individually rational under the Dual Error, establish sufficient conditions for center rationality, and derive an explicit welfare-loss bound. We further identify robustness effects: opposite-signed errors can cancel, vanishing-tail link functions (e.g., logistic) stabilize clear-cut cases via saturation, and extra noise smooths belief maps, reducing the gains from marginal manipulation. Experiments on simulations and real-world benchmarks show that EA-RAM is robust to the Dual Error and achieves a better cost--performance Pareto frontier than centralized baselines, with additional gains when providers contribute local information, validating its practical effectiveness.
Sources
- A Survey of Large Language Models
- ICL-Router: In-Context Learned Model Representations for LLM Routing
- COALESCE: Economic and Security Dynamics of Skill-Based Task Outsourcing Among Team of Autonomous LLM Agents
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Program Synthesis with Large Language Models
- Training Verifiers to Solve Math Word Problems
Related papers
- Exact Regret Frontiers and Externality Scheduling in Centralized Serial-Dictatorship Bandits
- In-Context Credit Assignment via the Core
- Breaking 1/epsilon Barrier in Quantum Zero-Sum Games: Generalizing Metric Subregularity for Spectraplexes
- Enhancing Affine Maximizer Auctions with Correlation-Aware Payment
- LLM Bidders Preserve the Mechanism-Level Orderings of Human Bidders
- Towards Performatively Stable Equilibria in Decision-Dependent Games for Arbitrary Data Distribution Maps