Error-Aware Reverse Auction Mechanism for Large Language Model Routing

arXiv:2608.12719 · cs.GT, cs.AI · Submitted 2026-08-13 · Read on arXiv

Shenzhen International Center for Industrial and Applied Mathematics · Shenzhen Research Institute of Big Data · The Chinese University of Hong Kong, Shenzhen · Shenzhen Campus of Sun Yat-sen University · Shenzhen Loop Area Institute

cs.GT, cs.AI

Submitted: 2026-08-13

Updated: 2026-09-28

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

Terminology

Summary

arXiv: 2608.12719v1 [cs.GT] 13 Aug 2026


The paper addresses the challenge of routing each query to a cost-effective large language model (LLM) to balance quality and cost. The authors identify two structural limitations of existing centralized routing paradigms:

  1. Information–risk mismatch: the task center bears the risk of failure but has less information about the LLMs, whereas providers hold richer private knowledge.

  2. Scalability bottleneck: the center must profile or train for every new model, incurring per-model overhead even for training-free routers, which hinders rapid expansion.

The authors propose a market-based routing paradigm that transfers ex-ante prediction to LLM providers via a reverse auction (Figure 1). In this paradigm:

  • The task center acts as the buyer and solicits bids from providers, who report their self-predicted success probabilities and costs

  • the center maintains only a model-agnostic ex-post evaluator

  • This distributed design aligns information with risk and removes the need for model-by-model profiling for the buyer, thereby improving scalability

The key theoretical contribution is explicitly modeling what the authors call the inherent Dual Error:

  • Providers' ex-ante success estimates are subjective and noisy

  • The center's ex-post evaluation is imperfect

Formally, the paper defines:

  • Ex-post evaluation error: The buyer evaluates the output via the error-involved ex-post acceptance probability hi = σ(ϕi + εpost), where the evaluation error εpost has E[εpost] = apost and Var(εpost) = bpost

  • Ex-ante prediction error: seller i predicts the buyer's acceptance probability... We introduce an independent prediction error εante,i with E[εante,i] = aante,i and Var(εante,i) = bante,i, and define the aggregated error ηi = εpost + εante,i

The seller's ex-ante subjective probability is gi = σ(ϕi + ηi), where E[ηi] = aη,i = apost + aante,i and Var(ηi) = bη,i = bpost + bante,i.

The mechanism works as follows (Algorithm 1):

  1. Bidding: Providers submit ex-ante bids with self-predicted success probabilities and costs

  2. Allocation: the center allocates the query based on reported surplus scores ŝi = V gi − ci

  3. Execution: The selected model answers

  4. Evaluation and Payment: The buyer evaluates the output and settles payment with payment rule rj = V µ̃j − H, where H is the runner-up score

The payment rule is designed so that payments depend on the evaluator signal µ̃j ∈ 0, 1, a noisy proxy for the unobserved ground truth µj rather than self-reported probabilities, preventing sellers from inflating their reported p̂i.

The paper establishes several key theoretical properties:

Theorem 3.2: Seller i's interim expected utility Uiseller(ŝi) is maximized at ŝi = T̄i. Hence, reporting ŝi = T̄i is a Bayesian best response.

Theorem 3.3: At the equilibrium ŝ = T̄, every seller satisfies IR: E[Uiseller θ̃i] ≥ 0.

Theorem 3.4: Two sufficient conditions guarantee CR:

  • (A) E[H] ≥ V∆gate, where ∆gate = Lσ√(bpost + a2post)

  • (B) ∆cons = E[p(1) − h(1)] ≥ 0

Theorem 3.7: The expected welfare loss satisfies 0 ≤ E[Wi⋆] − E[Wi†] = E[(V pi⋆−ci⋆) − (V pi†−ci†)] ≤ 2V Lσ(Mpost + Mante)

where Mpost = √(bpost + a2post) and Mante = maxi √(bante,i + a2ante,i).

The paper reveals three important robustness effects:

Proposition 3.8: If (gi − hi)(hi − pi) ≤ 0, for i ∈ i⋆, i†, then the welfare-loss bound tightens to 2V Lσ max Mante, Mpost. This means when ex-ante and ex-post errors have opposite signs, they partially cancel out, reducing the resulting welfare loss.

Proposition 3.9: For link functions with vanishing tails (e.g., logistic), the deviation ∆(ϕi) = E[σ(ϕi + ϵ)] − σ(ϕi) satisfies limϕi→∞ ∆(ϕi) = 0. This means flat-tailed links such as the logistic function saturate when ϕi is large, so score noise barely affects σ(ϕi) and the mechanism is robust in clear-cut cases.

Proposition 3.10: Injecting additional independent noise weakly reduces the maximal sensitivity of both ϕ ↦ gi(ϕ) and ϕ ↦ hi(ϕ). This limits the return to strategic manipulation.

The paper simulates N = 5 heterogeneous sellers with V = 20 and d = 1.5, using a highly competitive market with a small top-two surplus gap (∆s = 0.44), where mild noise can flip the ranking.

Results show:

  • Figure 5a: EA-RAM closely tracks this bound and maintains a stable welfare gap as σpost increases, whereas Error-Naive deteriorates sharply

  • Figure 5b: EA-RAM consistently outperforms Error-Naive, showing that internalizing the evaluation mechanism avoids severe misallocation and yields more graceful degradation

The paper evaluates on RouterBench, which aggregates per-query accuracy and monetary cost for 11 LLMs, spanning open-source (Llama, Mixtral) and proprietary models (GPT, Claude), across six benchmarks: HellaSwag, Winogrande, ARC-Challenge, MBPP, GSM8k, and MMLU.

Baselines compared: EmbedLLM, IRT-Router, RouteLLM, FrugalGPT, and Cascade Routing.

Key results:

  • Table 2: EA-RAM achieves the best AIQ (Average Improvement in Quality) across all benchmarks, with the base setting outperforming all baselines

  • Figure 6: EA-RAM already outperforms baselines in the base setting, and seller-side local information further shifts the frontier upward and leftward

  • Figure 7: In noisy LLM-as-a-Judge evaluation experiments, EA-RAM consistently performs best across all α, incurring the smallest welfare loss

  • Figure 8: The center-side latency of EA-RAM remains nearly constant as the model pool grows, since the center only performs evaluator and auction computation and does not require retraining or per-model prediction

  • Table 3: Communication overhead grows roughly linearly with N, while computation latency stays nearly constant, with communication becoming the bottleneck only below 0.338–0.384 MB/s per channel

The paper's main contributions are:

  1. "We propose EA-RAM, a market-based routing framework that shifts ex-ante prediction to LLM providers, addresses centralized routers' information–risk mismatch and per-model profiling bottleneck, and models LLM routing as an auction under explicit Dual Error."

  2. "We establish theoretical guarantees under Dual Error: EA-RAM satisfies BIC and IR, admits sufficient conditions for CR, and enjoys a welfare-loss bound; we further reveal robustness insights, including error compensation, saturation stability, and noise-induced flattening."

  3. Extensive experiments on simulations and real-world benchmarks show that EA-RAM is robust to Dual Error and improves economic efficiency compared to state-of-the-art centralized baselines.

The paper acknowledges several limitations:

  • Our present theory does not cover payment-rule variants such as bounded penalties or non-negative payments

  • in communication-heavy settings such as image- or video-based tasks, or under degraded network conditions, communication may become a more significant bottleneck

  • the auction protocol requires broadcasting each query to all candidate providers before selection, which may expose user prompts to non-winning providers and raises security concerns

  • providers may be incentivized to optimize toward the evaluator rather than the user's true underlying need

Future directions include: communication-efficient reverse-auction mechanisms, extending the theory to bounded-penalty or non-negative-payment variants, tightening the current welfare-loss bound, generalizing the framework beyond single-winner routing to support multi-provider cross-checking, cascading, and multi-turn settings, extending the model to heterogeneous risk preferences, and providing more principled support for non-binary evaluation tasks.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems:

  • Improvement: Replace centralized router prediction with a reverse auction where LLM providers bid with self-reported success probabilities and costs.

  • Capability: The AI system can dynamically select the most cost-effective model for each query without requiring the central system to profile or train on every new model. This enables seamless scaling as new LLMs are added to the pool.

  • Improvement: Explicitly model both ex-ante prediction errors (from providers) and ex-post evaluation errors (from the center) when making routing decisions.

  • Capability: The AI system can quantify uncertainty in both provider claims and its own evaluation, leading to more robust model selection that degrades gracefully under noisy conditions rather than failing catastrophically.

  • Improvement: Use payment rules based on noisy evaluator signals rather than self-reported probabilities, with payments tied to the runner-up score.

  • Capability: The AI system can incentivize truthful reporting from providers, preventing gaming of the routing system and ensuring that providers are rewarded based on actual performance rather than inflated claims.

  • Improvement: Detect when provider overestimation and center underestimation (or vice versa) partially cancel out, and adjust routing confidence accordingly.

  • Capability: The AI system can identify scenarios where errors naturally balance, allowing it to make higher-confidence routing decisions even when individual error sources are significant.

  • Improvement: Recognize that when model quality signals are extreme (very high or very low), noise has minimal impact on routing decisions.

  • Capability: The AI system can make rapid, confident routing decisions for queries where model suitability is obvious, while allocating more computational resources to ambiguous cases where noise could flip the ranking.

  • Improvement: Understand that adding controlled noise to evaluation signals reduces the sensitivity of outcomes to strategic manipulation.

  • Capability: The AI system can intentionally inject calibrated noise into its evaluation process to reduce the return on adversarial provider behavior, making the system more robust to gaming.

  • Improvement: Keep the central evaluation component independent of specific model architectures, focusing only on output quality.

  • Capability: The AI system can evaluate any new model's output without retraining, enabling instant integration of new LLMs into the routing pool with near-zero marginal cost.

  • Improvement: Balance the linear communication overhead of broadcasting queries to all providers against the constant computation latency of the center.

  • Capability: The AI system can operate efficiently in bandwidth-constrained environments by understanding when communication becomes the bottleneck (below 0.34 MB/s per channel) and adapting its routing strategy accordingly.

  • Improvement: Use the theoretical welfare-loss bound (≤ 2V Lσ(Mpost + Mante)) to guarantee that routing decisions remain economically efficient even with noisy predictions.

  • Capability: The AI system can provide formal guarantees on the maximum economic inefficiency introduced by prediction errors, enabling deployment in cost-sensitive enterprise applications where worst-case performance matters.

  • Improvement: Extend the single-winner auction framework to support cascading, cross-checking, and multi-turn interactions between multiple providers.

  • Capability: The AI system can orchestrate complex workflows where multiple models collaborate, verify each other's outputs, or iteratively refine responses, while maintaining the same incentive-compatibility guarantees.

Abstract

Routing each query to a cost-effective large language model (LLM) is critical for balancing quality and cost, yet most routers rely on a centralized task center to predict model performance, creating an information-risk mismatch and a scalability bottleneck as the model pool grows. We propose a market-based routing paradigm that shifts ex-ante prediction to LLM providers via a reverse auction, where providers bid with self-predicted success probabilities and execution costs. To account for inherently noisy provider predictions and center evaluations, we introduce the Error-Aware Reverse Auction Mechanism (EA-RAM), which explicitly models this inherent Dual Error. We prove that EA-RAM is Bayesian incentive compatible and individually rational under the Dual Error, establish sufficient conditions for center rationality, and derive an explicit welfare-loss bound. We further identify robustness effects: opposite-signed errors can cancel, vanishing-tail link functions (e.g., logistic) stabilize clear-cut cases via saturation, and extra noise smooths belief maps, reducing the gains from marginal manipulation. Experiments on simulations and real-world benchmarks show that EA-RAM is robust to the Dual Error and achieves a better cost--performance Pareto frontier than centralized baselines, with additional gains when providers contribute local information, validating its practical effectiveness.

Sources

Related papers