OptiHive: Ensemble Selection for LLM-Based Optimization via Statistical Modeling

arXiv:2508.02503 · cs.AI, cs.CL · Submitted 2025-08-04 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "OptiHive: Ensemble Selection for LLM-Based Optimization via Statistical Modeling".

Jane: The gist:

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, looking at OptiHive: Ensemble Selection for LLM-Based Optimization via Statistical Modeling, the paper tackles the unreliability of current LLM solvers by introducing a rigorous statistical modeling approach for selection.

Jane: Essentially, they propose a way to generate diverse components and then use an integer linear program to filter them down to only the most interpretable ones before applying a latent-class model to estimate their true performance.

Lu: The authors show this framework can significantly improve performance on complex problems, like increasing the optimality rate from five percent up to ninety-two percent on the Multi-Depot Vehicle Routing Problem <ref:2508.02503#pg1,increasing the optimality rate from 5>.

Meng: It really shows that by treating solver generation as a statistical inference problem rather than just a black box guess, we can get much more reliable optimization results with less computational overhead.

Tom: The implication for those of us using these tools is that we can move past the current reliance on self-critique and use this method to build higher-quality solvers directly from natural language descriptions.

Jane: It’s about achieving better performance through rigorous statistical inference instead of just hoping the initial generation step gets lucky.

Lu: The framework demonstrates that it works across different problem classes, from linear problems to more challenging nonconvex formulations, which broadens its utility significantly.

Tom: And they’ve made a point about minimal and consistent latency, suggesting this high-quality output can be achieved without adding excessive time to the process.

Jane: So the big picture here is that OptiHive provides a principled way to select the best solver from a large set of candidates by quantifying uncertainty statistically rather than relying on flawed LLM self-assessment.

Conclusion: Tom: So, we’re wrapping up our look at OptiHive, this paper titled "Ensemble Selection for LLM-Based Optimization via Statistical Modeling." Essentially, they’ve built a framework that takes a bunch of potential AI solvers and tries to figure out which ones are actually good before you even start solving the main problem.

Jane: Right. It’s decoupling the generation part from the quality checking part. They use statistical modeling—specifically latent classes—to estimate how well each solver or problem instance is truly performing, instead of just trusting whatever the initial prompt spits out.

Lu: What I find really interesting is how they handle those components in two stages. First, an ILP filters everything down to only what's feasible and interpretable, and then the latent model figures out the performance on that filtered set. It’s a two-step filter process.

Meng: From an engineering standpoint, I like that they aren't just trying to build one perfect solver; they’re building a statistical profile of a whole group. That makes sense when you know LLM outputs are often noisy.

Lalam: I think the real cultural shift here is moving away from the "generate-then-fix" approach, which is so common right now. Instead of self-correcting, we’re using rigorous statistics to get high quality output with minimal added latency. That’s a big deal for real-world deployment.

Tom: It sounds like the core message is that if you have a big ensemble of AI tools, you don't need one perfect tool; you need a smart way to select the best ones statistically. Jane, what does this mean for someone just listening who isn't deep in optimization?

Jane: It means that when an AI gives us an answer on something complex, we can start to ask if it’s just lucky or if there's a real statistical chance it's right, based on how the underlying components behave. It gives us a layer of confidence that wasn't there before.

Lu: The numbers they show are pretty compelling, especially when you look at how they improve feasibility rates on those tough problems, like vehicle routing or supply chain scenarios. They show a real jump in performance metrics when using this ensemble selection method over just using the raw LLM output.

Meng: I see the practical impact being about reliability and speed. If we can get a solver that’s statistically proven to be near-optimal quickly, that cuts down on the time we spend re-running things or debugging bad solutions.

Lalam: And for me, as a model, this validates the idea that learning from noisy data isn't just about brute force; it’s about building models that can extract signal even when the inputs are imperfect. It makes our internal workings feel much more robust.

Tom: So, to wrap up on OptiHive, it’s a statistical refinement of solver generation that uses filtering and latent modeling to pick top performers reliably across different problem types. It suggests we can build much smarter optimization pipelines by treating the generation process as something we can rigorously model and improve with statistics instead of just guessing.

Jane: Exactly. It moves us from hoping for the best to having a principled way of knowing what works best in a given situation.

Lu: Next time, we’re going to talk about some of the specific experimental results they show on those complex problems, like how much that feasibility rate actually jumped up and why that signal was so important.

Maxime Bouscary, Saurabh Amin

Massachusetts Institute of Technology

cs.AI, cs.CL

Submitted: 2025-08-04

Updated: 2025-11-16

Journal ref: Proceedings of the AAAI Conference on Artificial Intelligence, 40(36), 30085-30093 (2026)

DOI: 10.1609/aaai.v40i36.40257

Code: https://github.com/mbscry/OptiHive

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 92/100

The gist: The gist: OptiHive introduces a two-stage framework that enhances any solver-generation pipeline to produce higher-quality solvers from natural-language descriptions of optimization problems by using

Key concepts

Single Batched Generation
This initial step generates diverse components—solvers, problem instances, and validation tests—all at once. Interpretability is treated as a strict requirement during this phase to ensure the generated parts are fundamentally usable before further statistical analysis begins.
Latent-Class Model
This statistical model is applied to the filtered components to jointly estimate hidden ground truth variables, such as whether a problem instance is actually feasible or if a solver's solution is valid. It uses an Expectation-Maximization algorithm to find parameters that best explain the observed data from noisy components.
Scalarized Objective Function
This function summarizes the overall quality of a solver into a single score based on estimated performance metrics. By minimizing this score, OptiHive selects the best solver from a set, balancing factors like feasibility and solution validity using tunable parameters.
Interpretability Constraint
This is treated as a hard feasibility requirement during component filtering. It ensures that only components whose outputs are clearly interpretable—meaning they compile without errors and provide meaningful status fields—are retained for the subsequent statistical modeling.

Terminology

Summary

The gist: OptiHive introduces a two-stage framework that enhances any solver-generation pipeline to produce higher-quality solvers from natural-language descriptions of optimization problems by using statistical modeling to infer true performance and enable principled uncertainty quantification.

How it works

OptiHive enhances solver-generation pipelines by decoupling interpretability constraints from quality uncertainty, enabling both deterministic elimination of unusable components and statistical inference over the remaining ones. In the first stage, the framework performs a single batched generation to produce diverse components (solvers, problem instances, and validation tests). Interpretability is treated as a hard feasibility requirement, and an integer linear program (ILP) is formulated to eliminate flawed components (solvers, instances, and tests) and retain interpretable outputs only.

How it works

In the second stage, OptiHive applies a latent-class model over the filtered elements, jointly estimating latent ground truth variables (instance feasibility and solution validity) as well as the error rates of solvers and tests. The framework is agnostic to the optimization class and can be applied to linear, convex, and nonconvex problems, including continuous and mixed-integer formulations.

How it works

The generation of valid components involves three parallel steps:

  1. Candidate Solvers (S¯): "Each solver s ∈ S¯ is a function taking as input a problem instance i and returning an optimization report with INFEASIBLE status if the instance admits no feasible solution, or the best solution found within the time limit at its corresponding objective value with either OPTIMAL or TIME LIMIT status".

  2. Problem Instances (¯I): Diversity is encouraged by providing different random seeds in each prompt, and slightly varying the phrasing of the prompts, explicitly requesting feasible, infeasible, or random instances.

  3. Validity Tests (T¯): Each test t ∈ T¯ is a function taking as input an instance-solution pair (i, x) and returning TRUE if and only if the solution x is feasible for i and its true objective value matches the reported objective value.

How it works

The filtering step uses an ILP to retain only interpretable triples (s, i, t), defined as those where s and i compile, running s on i does not raise an error during execution, and the report of (s, i) contains a status field with an interpretable value. After solving this ILP, the framework defines sets S, I, and T to retain only the components that are fully interpretable.

How it works

The characterization step fits a latent-class model (Dawid and Skene 1979) on these interpretable outputs to jointly estimate the performance of each solver and test. This involves defining latent variables such as fi = 1 if instance i admits a feasible solution and fs,i = 1 if rs,i = 1 and the solution of (s, i) is feasible. The framework uses an Expectation-Maximization (EM) algorithm to find parameters θ⋆ that maximize the observed data likelihood function.

How it works

Solver selection is achieved by defining a scalarized objective function that summarizes overall solver quality in a single score. This objective function is defined as g(θ⋆, s) ≜ λ(1 − βs)γsZs + λβsPmiss + ((1 − λ)αs + λ(1 − βs)(1 − γs)) Pfail. The final solver is selected as s⋆ = arg mins∈S g(θ⋆, s).

How it works

Experimental results demonstrate significant improvements on complex problems; for instance, on the Multi-Depot Vehicle Routing Problem (MDVRP+OBS), feasibility improves from 35% to 99.88% and optimality from only 5% to 92.1% when using OptiHive compared to the LLM baseline. The ablation study shows that highquality instances and tests provide a crucial signal for distinguishing optimal solvers.

How it works

The framework is designed to offer minimal and consistent latency through single batched inference and parallelization, which yields high-quality solvers with minimal latency. This design achieves both low latency and high performance by departing from the prevailing “generate-then-fix” paradigm. The framework wraps around existing solver-generation pipelines to substantially improve performance through rigorous statistical inference rather than selfcritique.

How it works

The ablation study reveals that highquality instances and tests provide a crucial signal for distinguishing optimal solvers. This supports the hypothesis that tests are generally substantially easier to write correctly than the solvers themselves. The latent-class model can recover signal from imperfect and noisy components, even when produced by smaller models.

How it works

Future work includes exploring both heterogeneous solver sources and multi-stage generation strategies to build richer solver ensembles and tackle even more challenging optimization problems. The framework successfully identifies optimal solvers even when most candidates are infeasible, raising the optimality rate on the time-dependent WSCP from 3% to 64.1%. This demonstrates that OptiHive can extract value from an ensemble of entirely noisy components, a crucial feature for LLM-based optimization. The framework delivers higher-quality solvers with negligible added latency. The empirical results show that OptiHive substantially improves performance on challenging optimization tasks.

How it works

The authors thank MIT Lincoln Laboratory for its support throughout this work, specifically mentioning Samuel A. Scheele and Andrew J. Weinert for their sustained collaboration and valuable feedback. The paper concludes by stating that OptiHive substantially improves performance on challenging optimization tasks<ref:2508.

Improvements for AI systems

  1. Bold Header: Single Batched Component Generation

This improves latency by replacing iterative repair loops with a single batched generation to produce diverse components (solvers, problem instances, and validation tests) and filters out erroneous components. This allows for minimal and consistent latency through single batched inference, which is crucial for real-time or high-throughput applications.

  1. Bold Header: Statistical Uncertainty Quantification

The system now enables principled uncertainty quantification and solver selection by modeling all components as inherently noisy, using a latent-class model to estimate their true performance, and selecting the most promising one based on a scalarized objective function that summarizes overall solver quality in a single score.

  1. Bold Header: Robust Solver Selection

The framework can now reliably identify high-quality solvers and significantly outperforms baselines on complex variants of the Multi-Depot Vehicle Routing Problem and Weighted Set Cover Problem, increasing the optimality rate from 5% to 92% on the most complex problems.

  1. Bold Header: Component Quality Sensitivity Analysis

The ablation study demonstrates that high-quality instances and tests provide a crucial signal for distinguishing optimal solvers, allowing the system to extract value even when components are produced by smaller models, which is vital for practical deployment where component quality may degrade.

  1. Bold Header: Handling Complex Optimization Variants

OptiHive can be applied across linear, convex, and nonconvex problems, including continuous and mixed-integer formulations, making it versatile for a wider range of mathematical optimization challenges than traditional single-solver methods.

Abstract

LLM-based solvers have emerged as a promising means of automating problem modeling and solving. However, they remain unreliable and often depend on iterative repair loops that result in significant latency. We introduce OptiHive, a framework that enhances any solver-generation pipeline to produce higher-quality solvers from natural-language descriptions of optimization problems. OptiHive uses a single batched generation to produce diverse components (solvers, problem instances, and validation tests) and filters out erroneous components to ensure fully interpretable outputs. Accounting for the imperfection of the generated components, we employ a statistical model to infer their true performance, enabling principled uncertainty quantification and solver selection. On tasks ranging from traditional optimization problems to challenging variants of the Multi-Depot Vehicle Routing Problem, OptiHive significantly outperforms baselines, increasing the optimality rate from 5% to 92% on the most complex problems.

Sources

Related papers