Online Pandora's Box for Contextual LLM Cascading
cs.AI, cs.LG, econ.EM, stat.ML
Submitted: 2026-06-05
Updated: 2026-08-26
License: http://creativecommons.org/licenses/by/4.0/
The gist: Motivated by Large Language Model (LLM) cascading, we propose an online contextual Pandora's Box model for adaptively querying and selecting LLM APIs.
Terminology
Abstract
Motivated by Large Language Model (LLM) cascading, we propose an online contextual Pandora's Box model for adaptively querying and selecting LLM APIs. In each period, a decision-maker observes a request context and faces a two-phase decision problem. In the query phase, the decision-maker sequentially queries APIs, where each query reveals a generated output and the decision-maker incurs an (output-dependent) cost. In the selection phase, the decision-maker selects one of the generated outputs to deploy and observes only the downstream reward of the deployed output. This output-mediated feedback structure differs from classical online contextual Pandora's Box models, in which opening a box directly reveals its reward. Rather than estimating the full conditional output and cost distributions of each API, we directly model the reservation index and develop a learning approach for the query phase. Specifically, we impose a parametric structure on the contextual reservation index functions induced by the classical Weitzman's policy. Our policy combines generalized method of moments (GMM) type estimation of these reservation indices with UCB-style confidence bounds for both these indices and the shared output-level reward evaluator. Under regularity conditions, we prove that the resulting policy achieves dimension-dependent O(sqrt T) cumulative regret over a horizon of T periods.
Sources
- Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints
- Language Model Cascades: Token-level uncertainty and beyond
- Tryage: Real-time, intelligent Routing of User Prompts to Large Language Models
- RouterBench: A Benchmark for Multi-LLM Routing System
- Optimal Bayesian Stopping for Efficient Inference of Consistent LLM Answers
- Online Scheduling for LLM Inference with KV Cache Constraints
- LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion
- Asymptotically Optimal Sequential Testing with Heterogeneous LLMs
- Improved Regret and Contextual Linear Extension for Pandora's Box and Prophet Inequality
- Online Cascade Learning for Efficient Inference over Streams
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection