Frontier LLMs are effective batch optimizers: Assessing reasoning models in continuous and discrete settings
cs.LG
Submitted: 2026-09-02
Updated: 2026-09-02
License: http://creativecommons.org/licenses/by/4.0/
The gist: Frontier large language models (LLMs) have become attractive priors for optimization due to their large-scale pretraining that enables them to navigate a variety of optimization settings.
Terminology
Abstract
Frontier large language models (LLMs) have become attractive priors for optimization due to their large-scale pretraining that enables them to navigate a variety of optimization settings. However, the effectiveness of modern reasoning LLMs in batch optimization settings remains underexplored. Here we investigate the performance of the current generation of frontier LLMs as batch optimizers in both continuous and discrete settings. We find that while LLMs are competitive zero-shot batch optimizers for numerical test functions, their performance is brittle compared to classical non-LLM optimization approaches. However, LLM priors are significantly better in semantically rich settings, indicating that their batch optimization behavior is highly effective when navigating and reasoning over the discrete spaces most similar in structure to their pretraining data.
Sources
- Composer 2 Technical Report
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Evaluating the Progression of Large Language Model Capabilities for Small-Molecule Drug Design
- Training a Scientific Reasoning Model for Chemistry
- Reinforced In-Context Black-Box Optimization
- ZeroShotOpt: Towards Zero-Shot Pretrained Models for Efficient Black-Box Optimization
- Large Language Models to Enhance Bayesian Optimization
- Large Language Models as Optimizers
- Sample Efficiency Matters: A Benchmark for Practical Molecular Optimization
- A Tutorial on Bayesian Optimization
- Practical Bayesian Optimization of Machine Learning Algorithms
- BioBO: Biology-informed Bayesian Optimization for Perturbation Design
- Bayesian Optimization of Catalysis With In-Context Learning
- Exploring the True Potential: Evaluating the Black-box Optimization Capability of Large Language Models
- GP-MoLFormer-Sim: Test Time Molecular Optimization through Contextual Similarity Guidance
- Sample Efficient Generative Optimization for Molecular Design
- ChemBERTa-2: Towards Chemical Foundation Models
- ExLLM: Experience-Enhanced LLM Optimization for Molecular Design and Beyond
- Efficient Evolutionary Search Over Chemical Space with Large Language Models
- GenMol: A Drug Discovery Generalist with Discrete Diffusion
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks