Online Learning with LLM Experts from Limited Feedback
cs.LG, stat.ML
Submitted: 2026-09-05
Updated: 2026-09-05
Comments: 21 pages, 4 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: We study adaptive routing of prompts to large language model (LLM) experts to maximize response quality in an online setting with limited feedback.
Terminology
Abstract
We study adaptive routing of prompts to large language model (LLM) experts to maximize response quality in an online setting with limited feedback. We formulate it as a bandit problem with K actions that represent experts and d features that encode prompts, over a horizon of T rounds. We propose algorithms that strategically select and observe rewards to minimize regret. In the full-information setting, we achieve a regret of (d T / sqrt m), while in the bandit setting we achieve (d T sqrt K / m), where m T is a budget on feedback. Our experiments show that we efficiently learn high-quality routing strategies across diverse LLMs from limited feedback.
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks