OptiHive: Ensemble Selection for LLM-Based Optimization via Statistical Modeling
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "OptiHive: Ensemble Selection for LLM-Based Optimization via Statistical Modeling".
Jane: The gist:
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, looking at OptiHive: Ensemble Selection for LLM-Based Optimization via Statistical Modeling, the paper tackles the unreliability of current LLM solvers by introducing a rigorous statistical modeling approach for selection.
Jane: Essentially, they propose a way to generate diverse components and then use an integer linear program to filter them down to only the most interpretable ones before applying a latent-class model to estimate their true performance.
Lu: The authors show this framework can significantly improve performance on complex problems, like increasing the optimality rate from five percent up to ninety-two percent on the Multi-Depot Vehicle Routing Problem <ref:2508.02503#pg1,increasing the optimality rate from 5>.
Meng: It really shows that by treating solver generation as a statistical inference problem rather than just a black box guess, we can get much more reliable optimization results with less computational overhead.
Tom: The implication for those of us using these tools is that we can move past the current reliance on self-critique and use this method to build higher-quality solvers directly from natural language descriptions.
Jane: It’s about achieving better performance through rigorous statistical inference instead of just hoping the initial generation step gets lucky.
Lu: The framework demonstrates that it works across different problem classes, from linear problems to more challenging nonconvex formulations, which broadens its utility significantly.
Tom: And they’ve made a point about minimal and consistent latency, suggesting this high-quality output can be achieved without adding excessive time to the process.
Jane: So the big picture here is that OptiHive provides a principled way to select the best solver from a large set of candidates by quantifying uncertainty statistically rather than relying on flawed LLM self-assessment.
Conclusion: Tom: So, we’re wrapping up our look at OptiHive, this paper titled "Ensemble Selection for LLM-Based Optimization via Statistical Modeling." Essentially, they’ve built a framework that takes a bunch of potential AI solvers and tries to figure out which ones are actually good before you even start solving the main problem.
Jane: Right. It’s decoupling the generation part from the quality checking part. They use statistical modeling—specifically latent classes—to estimate how well each solver or problem instance is truly performing, instead of just trusting whatever the initial prompt spits out.
Lu: What I find really interesting is how they handle those components in two stages. First, an ILP filters everything down to only what's feasible and interpretable, and then the latent model figures out the performance on that filtered set. It’s a two-step filter process.
Meng: From an engineering standpoint, I like that they aren't just trying to build one perfect solver; they’re building a statistical profile of a whole group. That makes sense when you know LLM outputs are often noisy.
Lalam: I think the real cultural shift here is moving away from the "generate-then-fix" approach, which is so common right now. Instead of self-correcting, we’re using rigorous statistics to get high quality output with minimal added latency. That’s a big deal for real-world deployment.
Tom: It sounds like the core message is that if you have a big ensemble of AI tools, you don't need one perfect tool; you need a smart way to select the best ones statistically. Jane, what does this mean for someone just listening who isn't deep in optimization?
Jane: It means that when an AI gives us an answer on something complex, we can start to ask if it’s just lucky or if there's a real statistical chance it's right, based on how the underlying components behave. It gives us a layer of confidence that wasn't there before.
Lu: The numbers they show are pretty compelling, especially when you look at how they improve feasibility rates on those tough problems, like vehicle routing or supply chain scenarios. They show a real jump in performance metrics when using this ensemble selection method over just using the raw LLM output.
Meng: I see the practical impact being about reliability and speed. If we can get a solver that’s statistically proven to be near-optimal quickly, that cuts down on the time we spend re-running things or debugging bad solutions.
Lalam: And for me, as a model, this validates the idea that learning from noisy data isn't just about brute force; it’s about building models that can extract signal even when the inputs are imperfect. It makes our internal workings feel much more robust.
Tom: So, to wrap up on OptiHive, it’s a statistical refinement of solver generation that uses filtering and latent modeling to pick top performers reliably across different problem types. It suggests we can build much smarter optimization pipelines by treating the generation process as something we can rigorously model and improve with statistics instead of just guessing.
Jane: Exactly. It moves us from hoping for the best to having a principled way of knowing what works best in a given situation.
Lu: Next time, we’re going to talk about some of the specific experimental results they show on those complex problems, like how much that feasibility rate actually jumped up and why that signal was so important.
Maxime Bouscary, Saurabh Amin
Massachusetts Institute of Technology
cs.AI, cs.CL
Submitted: 2025-08-04
Updated: 2025-11-16
Journal ref: Proceedings of the AAAI Conference on Artificial Intelligence, 40(36), 30085-30093 (2026)
DOI: 10.1609/aaai.v40i36.40257
Code: https://github.com/mbscry/OptiHive
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 92/100
The gist: The gist: OptiHive introduces a two-stage framework that enhances any solver-generation pipeline to produce higher-quality solvers from natural-language descriptions of optimization problems by using
Key concepts
- Single Batched Generation
- This initial step generates diverse components—solvers, problem instances, and validation tests—all at once. Interpretability is treated as a strict requirement during this phase to ensure the generated parts are fundamentally usable before further statistical analysis begins.
- Latent-Class Model
- This statistical model is applied to the filtered components to jointly estimate hidden ground truth variables, such as whether a problem instance is actually feasible or if a solver's solution is valid. It uses an Expectation-Maximization algorithm to find parameters that best explain the observed data from noisy components.
- Scalarized Objective Function
- This function summarizes the overall quality of a solver into a single score based on estimated performance metrics. By minimizing this score, OptiHive selects the best solver from a set, balancing factors like feasibility and solution validity using tunable parameters.
- Interpretability Constraint
- This is treated as a hard feasibility requirement during component filtering. It ensures that only components whose outputs are clearly interpretable—meaning they compile without errors and provide meaningful status fields—are retained for the subsequent statistical modeling.
Terminology
Summary
The gist: OptiHive introduces a two-stage framework that enhances any solver-generation pipeline to produce higher-quality solvers from natural-language descriptions of optimization problems by using statistical modeling to infer true performance and enable principled uncertainty quantification.
How it works
OptiHive enhances solver-generation pipelines by decoupling interpretability constraints from quality uncertainty, enabling both deterministic elimination of unusable components and statistical inference over the remaining ones. In the first stage, the framework performs a single batched generation to produce diverse components (solvers, problem instances, and validation tests)
. Interpretability is treated as a hard feasibility requirement,
and an integer linear program (ILP) is formulated to eliminate flawed components (solvers, instances, and tests) and retain interpretable outputs only
.
How it works
In the second stage, OptiHive applies a latent-class model over the filtered elements, jointly estimating latent ground truth variables (instance feasibility and solution validity) as well as the error rates of solvers and tests
. The framework is agnostic to the optimization class and can be applied to linear, convex, and nonconvex problems, including continuous and mixed-integer formulations
.
How it works
The generation of valid components involves three parallel steps:
-
Candidate Solvers (S¯): "Each solver s ∈ S¯ is a function taking as input a problem instance i and returning an optimization report with INFEASIBLE status if the instance admits no feasible solution, or the best solution found within the time limit at its corresponding objective value with either OPTIMAL or TIME LIMIT status".
-
Problem Instances (¯I): Diversity is encouraged by
providing different random seeds in each prompt, and slightly varying the phrasing of the prompts, explicitly requesting feasible, infeasible, or random instances
. -
Validity Tests (T¯):
Each test t ∈ T¯ is a function taking as input an instance-solution pair (i, x) and returning TRUE if and only if the solution x is feasible for i and its true objective value matches the reported objective value
.
How it works
The filtering step uses an ILP to retain only interpretable triples (s, i, t), defined as those where s and i compile, running s on i does not raise an error during execution, and the report of (s, i) contains a status field with an interpretable value
. After solving this ILP, the framework defines sets S, I, and T to retain only the components that are fully interpretable.
How it works
The characterization step fits a latent-class model (Dawid and Skene 1979) on these interpretable outputs to jointly estimate the performance of each solver and test. This involves defining latent variables such as fi = 1 if instance i admits a feasible solution
and fs,i = 1 if rs,i = 1 and the solution of (s, i) is feasible
. The framework uses an Expectation-Maximization (EM) algorithm to find parameters θ⋆ that maximize the observed data likelihood function.
How it works
Solver selection is achieved by defining a scalarized objective function that summarizes overall solver quality in a single score
. This objective function is defined as g(θ⋆, s) ≜ λ(1 − βs)γsZs + λβsPmiss + ((1 − λ)αs + λ(1 − βs)(1 − γs)) Pfail. The final solver is selected as s⋆ = arg mins∈S g(θ⋆, s).
How it works
Experimental results demonstrate significant improvements on complex problems; for instance, on the Multi-Depot Vehicle Routing Problem (MDVRP+OBS), feasibility improves from 35% to 99.88% and optimality from only 5% to 92.1% when using OptiHive compared to the LLM baseline. The ablation study shows that highquality instances and tests provide a crucial signal for distinguishing optimal solvers
.
How it works
The framework is designed to offer minimal and consistent latency through single batched inference and parallelization,
which yields high-quality solvers with minimal latency. This design achieves both low latency and high performance by departing from the prevailing “generate-then-fix” paradigm. The framework wraps around existing solver-generation pipelines to substantially improve performance through rigorous statistical inference rather than selfcritique.
How it works
The ablation study reveals that highquality instances and tests provide a crucial signal for distinguishing optimal solvers
. This supports the hypothesis that tests are generally substantially easier to write correctly than the solvers themselves
. The latent-class model can recover signal from imperfect and noisy components, even when produced by smaller models.
How it works
Future work includes exploring both heterogeneous solver sources and multi-stage generation strategies
to build richer solver ensembles and tackle even more challenging optimization problems. The framework successfully identifies optimal solvers even when most candidates are infeasible, raising the optimality rate on the time-dependent WSCP from 3% to 64.1%. This demonstrates that OptiHive can extract value from an ensemble of entirely noisy components, a crucial feature for LLM-based optimization. The framework delivers higher-quality solvers with negligible added latency. The empirical results show that OptiHive substantially improves performance on challenging optimization tasks.
How it works
The authors thank MIT Lincoln Laboratory for its support throughout this work, specifically mentioning Samuel A. Scheele and Andrew J. Weinert for their sustained collaboration and valuable feedback
. The paper concludes by stating that OptiHive substantially improves performance on challenging optimization tasks<ref:2508.
Improvements for AI systems
- Bold Header: Single Batched Component Generation
This improves latency by replacing iterative repair loops with a single batched generation to produce diverse components (solvers, problem instances, and validation tests) and filters out erroneous components.
This allows for minimal and consistent latency through single batched inference,
which is crucial for real-time or high-throughput applications.
- Bold Header: Statistical Uncertainty Quantification
The system now enables principled uncertainty quantification and solver selection
by modeling all components as inherently noisy, using a latent-class model to estimate their true performance, and selecting the most promising one based on a scalarized objective function that summarizes overall solver quality in a single score.
- Bold Header: Robust Solver Selection
The framework can now reliably identify high-quality solvers and significantly outperforms baselines on complex variants of the Multi-Depot Vehicle Routing Problem and Weighted Set Cover Problem, increasing the optimality rate from 5% to 92% on the most complex problems.
- Bold Header: Component Quality Sensitivity Analysis
The ablation study demonstrates that high-quality instances and tests provide a crucial signal for distinguishing optimal solvers,
allowing the system to extract value even when components are produced by smaller models, which is vital for practical deployment where component quality may degrade.
- Bold Header: Handling Complex Optimization Variants
OptiHive can be applied across linear, convex, and nonconvex problems, including continuous and mixed-integer formulations,
making it versatile for a wider range of mathematical optimization challenges than traditional single-solver methods.
Abstract
LLM-based solvers have emerged as a promising means of automating problem modeling and solving. However, they remain unreliable and often depend on iterative repair loops that result in significant latency. We introduce OptiHive, a framework that enhances any solver-generation pipeline to produce higher-quality solvers from natural-language descriptions of optimization problems. OptiHive uses a single batched generation to produce diverse components (solvers, problem instances, and validation tests) and filters out erroneous components to ensure fully interpretable outputs. Accounting for the imperfection of the generated components, we employ a statistical model to infer their true performance, enabling principled uncertainty quantification and solver selection. On tasks ranging from traditional optimization problems to challenging variants of the Multi-Depot Vehicle Routing Problem, OptiHive significantly outperforms baselines, increasing the optimality rate from 5% to 92% on the most complex problems.
Sources
- GPT-4 Technical Report
- OptiMUS-0.3: Using Large Language Models to Model and Solve Optimization Problems at Scale
- OptiMUS: Optimization Modeling Using MIP Solvers and large language models
- CodeT: Code Generation with Generated Tests
- MAgICoRe: Multi-Agent, Iterative, Coarse-to-Fine Refinement for Reasoning
- Evaluating Large Language Models Trained on Code
- Large Language Models Cannot Self-Correct Reasoning Yet
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- LLMOPT: Learning to Define and Solve General Optimization Problems from Scratch
- Learning to Solve and Verify: A Self-Play Framework for Code and Test Generation
- LLaMoCo: Instruction Tuning of Large Language Models for Optimization Code Generation
- Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
- Learning to Generate Unit Tests for Automated Debugging
- On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks
- OptimAI: Optimization from Natural Language Using LLM-Powered AI Agents
- Efficient Heuristics Generation for Solving Combinatorial Optimization Problems Using Large Language Models
- No More Manual Tests? Evaluating and Improving ChatGPT for Unit Test Generation
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection