Bandits in Prod: Hyperparameter Optimization at Inference Time

arXiv:2609.01335 · cs.LG, cs.AI · Submitted 2026-09-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Bandits in Prod: Hyperparameter Optimization at Inference Time".

Jane: The paper was written by Louis Abraham, Tuan-Anh Nguyen and Nicolas Devatine from Tiime.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Tom: The paper formalizes this entire process by framing it as an infinitely many-armed bandit, or IMAB, where each potential configuration is a separate arm of the bandit.

Jane: Since there are so many possible combinations—temperature choices combined with model selections—the number of arms vastly outnumbers the requests we will ever send to them in production.

Lu: It’s not just a standard multi-armed bandit because, as Tom mentioned, it’s an infinite space; you can't just pull from a fixed list and wait for is to end.

Meng: The authors propose "IMABO," which manages this problem by combining a bandit policy that selects known configurations with an oracle that suggests entirely new ones to be added.

Lalam: That distinction is key because it allows us to continuously explore the vast search space without needing some predefined, pre-existing menu of options.

Tom: And the core idea is anchored in IMOSS, which provides a sophisticated anytime bandit policy designed specifically for this ongoing process.

Improvements: Jane: The paper really shines in how it addresses the specific choices we make about how to find new arms, moving beyond just picking them randomly.

Tom: They introduce four specific oracles that serve as different ways to generate those promising new configurations, like IMOSS-TPE and IMOSS-TabPFN.

Lu: I'm particularly interested in how they are applying concepts from classical optimization theory to these dynamic, live settings where the structure is so complex.

Meng: The practical benefit of using an oracle instead of random sampling is massive; we are actively steering the search toward regions that have shown promise based on past rewards.

Lalam: It feels like a much smarter way to manage resources, knowing that if a model performs well, we want to keep exploring variations around it rather than starting over from scratch.

Tom: The paper shows that these learned oracles consistently outperform the simple uniform random baseline across both classical machine-learning tasks and more complex LLM-based agents.

Conclusion: Jane: We've seen how this framework works on various benchmarks, from discrete classification problems to continuous optimization of things like learning rates.

Tom: The results are incredibly robust, suggesting that "Bandits in Prod" is a generalized solution for whatever online hyperparameter optimization problem you face.

Lu: The theoretical result regarding the cumulative rho-regret bound is also a huge accomplishment, giving us mathematical guarantees about how quickly we will find an optimal region of the search space.

Meng: We found that even when dealing with messy, non-separable problems like Gaussian functions, or "HotpotQA," the approach maintains its efficiency and performance.

Lalam: It gives me great hope for future systems because it suggests that we don't have to settle for a good configuration; we can continuously refine it over time.

Tom: That’s exactly what they found—the system is designed to be constantly improving, rather than just finding the best static choice.

Final Wrap-up: Tom: So, after seeing the evidence across all benchmarks, it seems clear that "Bandits in Prod: Hyperparameter Optimization at Inference Time" offers a powerful solution for managing live AI tuning.

Jane: It’ provides a framework that is both theoretically sound and practically effective for continuous improvement.

Lu: I think the power of this work lies in its ability to capture the dynamic nature of finding an optimal configuration, moving beyond static assumptions about the search space.

Meng: For us in industry, it means we can deploy more complex models with higher confidence that our tuning process is actively working to achieve peak performance.

Lalam: I believe this methodology will fundamentally change how we perceive quality in AI, prioritizing continuous learning over a singular static achievement.

Tom: We certainly have a lot of excitement about this research going into the world, and I think it’s been really helpful to break down all those complex findings for our listeners today.

Louis Abraham, Tuan-Anh Nguyen, Nicolas Devatine

Tiime

cs.LG, cs.AI

Submitted: 2026-09-01

Updated: 2026-09-02

Code: https://github.com/Tiime-Software/IMABO

Importance score: 93/100

The gist: The paper explores advanced methods for hyperparameter optimization, particularly focusing on how these techniques function when deployed in a production or inference environment.

Key concepts

Infinitely Many-Armed Bandit (IMAB)
This framework formalizes the optimization process. Each potential configuration, such as a specific temperature choice or model selection, is treated as a separate 'arm' of the bandit. Because there are so many combinations, this search space is considered infinite.
IMABO
This proposed mechanism manages the infinite search space. It combines a standard bandit policy that selects known configurations with an oracle. The oracle suggests entirely new, promising configurations that were not previously part of the fixed options.

Terminology

Summary

The paper explores advanced methods for hyperparameter optimization, particularly focusing on how these techniques function when deployed in a production or inference environment. It details various strategies—such as local versus global proposals—and analyzes the underlying cost structures of exploring different regions of the search space. Understanding these mechanisms is critical because effective optimization at inference time can determine the performance and reliability of complex AI systems, especially those dealing with high-stakes applications like question answering.

Local vs. Global Proposal Strategies

The core difference between local and global proposals lies in how they guide the policy's exploration. A local oracle proposes arms that are one parameter away from a good configuration found so far, causing the policy to wander among arms that are all nearly as good as the best one. This approach means that while the local oracle does not make the policy a better chooser, it effectively makes the choice matter less by keeping exploration confined to promising areas. Conversely, a global proposal keeps adding arms from all over the space, where such wandering is significantly more expensive.

However, locality introduces a potential weakness: a one-parameter change cannot take the search far from where it already is. This limitation proves problematic in certain tasks; for instance, on Higgs, the good region needs several parameters set correctly at once, meaning that single-step local movements are insufficient to reach the optimal configuration. This gap is the specific scenario where global methods, such as TabPFN, which models the whole reward surface, demonstrate superiority by admitting the best arm of the four oracles.

Analyzing Exploration Cost and Failure

The cost incurred when a pull lands on an arm that is not the best admitted so far is analyzed by factoring it into two quantities: how often a pull lands somewhere other than the best arm admitted so far, and how much it costs when it does. Table 4 illustrates that for nearly every benchmark and oracle, the policy is off the best arm for the great majority of pulls.

The data shows that local oracles are particularly effective in minimizing loss when deviations occur. For example, under the TPE oracle, the mean loss per pull not on the best arm ranges from 0.020 to 0.057, while under TabPFN, it is often lower (e.g., 0.017 for Mixed spaces). The analysis concludes that the two local ones [oracles]... lose the least when it happens, suggesting that confining the search space minimizes the penalty associated with suboptimal exploration.

Practical Guidance and Optimization Implementation

The authors provide specific recommendations for practical implementation, advising users to focus on mutating the current best configuration. They suggest resorting to a surrogate only when the optimum is expected to be difficult to reach. On many tasks, the model-free local oracle match[es] the transformer-based one while fitting a single one-dimensional density per proposal rather than performing a forward pass through a pretrained network.

For baseline comparison, the global TPE proposal is described as the metric-free baseline that the mutation oracle improves on. The global TPE model does not require metrics on the space; instead, it models only the density ratio between promising and unpromising configurations rather than the full reward surface, allowing it to propose over a continuum without needing a pre-declared grid.

HotpotQA RAG Pipeline Details

The experimental setup for HotpotQA utilizes a Retrieval-Augmented Generation (RAG) pipeline where the search space is defined by several hyperparameters. The system uses three prompt templates—few shot, zero shot, and naive—which dictate how the model processes context and questions. At inference time, the selected template is sent as the system message, while the user message combines the top-k retrieved passages and the HotpotQA query. The system also incorporates robust failure handling: Failed calls are retried up to 20 times with random exponential backoff for common service errors.

Improvements for AI systems

The existing work provides significant empirical guidance on optimizing search strategies by comparing local mutation versus global proposals. However, the current methodology relies on distinct oracles and treats the cost function (Cost(pull best arm)) as a secondary, descriptive metric rather than an integrated component of the acquisition function.

We propose developing a Hybrid Adaptive Optimization Engine (HAO) that dynamically selects and weights search strategies based on an estimated local manifold structure and global optimization difficulty.

The primary improvement is to formalize the trade-off between expected reward gain (Exploitation/Improvement) and the cost of perturbation (Exploration/Wandering). We must move beyond merely measuring mean loss per pull and incorporate this cost directly into the acquisition function.

Mechanism:

We will modify standard Bayesian Optimization acquisition functions (e.g., Expected Improvement, EI) to adopt a weighted Hamiltonian-like dynamic:

A hybrid(x) = E[(0, f(x) - f best)] - lambda times Cost(x mu,)

  • Term 1 (Expected Improvement): Standard exploitation term.

  • Term 2 (Cost Penalty): This is the critical addition. Cost(x mu,) estimates the expected cost associated with proposing a point x given the current local model parameters (mu,). This cost estimate must be derived from the local density ratio and projected onto a learned manifold.

  • ** lambda (Adaptivity Weight):** A task-specific hyperparameter that controls the relative importance of exploration versus exploitation.

The current system treats local search as simple mutation, which fails when the optimal region is highly constrained and requires multiple parameters to be set correctly simultaneously (the Higgs case). We must detect when simple one-parameter steps are insufficient.

  1. If Rank(Hessian) about 1 and Eigenvalues 1: The search space is locally smooth and dominated by single-parameter movement. The system defaults to the efficient Local Mutation Oracle.

  2. If Rank(Hessian) > 2 and Eigenvalues are large/mixed: This indicates a highly constrained, multi-dimensional optimal manifold (like the Higgs example). The system must switch to a Constrained Global Proposal Generator, which uses the principles of TabPFN—modeling the relationship between parameters in discrete or semi-discrete subspaces—but integrates this model into the gradient calculation for continuous variables.

  3. If Local Gradient Variance is high: This suggests the local region is noisy or highly non-stationary. The system temporarily reverts to a Metric-Free Global Baseline (TPE/Continual Search) to escape potential local minima, but with a reduced exploration radius guided by the cost penalty term (lambda).

Instead of relying on one oracle at a time, we propose integrating the strengths of all models into a hierarchical framework.

The final acquisition decision is then a weighted vote: A hybrid = w L times A(S L) + w G times A(S G) + w M times Guidance(M opt). The weights (w) are dynamically determined by the LCE and the observed variance of the reward function, providing a robust blend that maximizes both immediate gain and global exploration potential.


The resulting Hybrid Adaptive Optimization Engine (HAO) can perform:

  1. Adaptive Search Pathing: It will autonomously determine whether the optimal search path requires minute, local adjustments (e.g., fine-tuning a single hyperparameter by 10-4), or if it requires a significant, multi-dimensional jump across the parameter space (e.g., finding an optimal combination

Sources

Related papers