Pretrained Optimization Model for Zero-Shot Black Box Optimization
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Pretrained Optimization Model for Zero-Shot Black Box Optimization".
Jane: The paper was written by Xiaobin Li, Yujian Betterrest Li, Kai Wu, Xiaoyu Zhang, Handing Wang et al. from Xidian University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everybody. Today we are digging into a brand new paper that just hit arXiv, and it’s called “Pretrained Optimization Model for Zero-Shot Black Box Optimization.” Jane, I have to say, that title alone got me excited.
Jane: Same here, Tom. And honestly, the title is doing a lot of work. Let’s break it down. “Black box optimization” means you have a problem where you can test solutions and get a score, but you have no idea what’s happening inside. No gradients, no formulas. You just try something, see how good it is, and try again.
Tom: Right, and that’s how a lot of real-world tuning works. You’re tweaking a neural network’s hyperparameters, or you’re trying to control a robot, and you only get a reward signal. You don’t get a nice clean equation to solve.
Jane: Exactly. And the “zero-shot” part is the real kicker. That means you take this optimizer, you’ve never seen this specific problem before, and you just apply it directly. No retraining, no fiddling with settings. You just drop it in and let it run.
Tom: And that’s the dream, right? Because most optimizers are finicky. You change the problem slightly, and suddenly you have to spend hours tuning the algorithm’s own knobs. This paper wants to kill that whole workflow.
Jane: The authors are from Xidian University, and they’ve built something they call POM. It’s a pretrained model, like a language model, but instead of learning to predict the next word, it learns to generate better solutions to optimization problems.
Tom: So instead of training a model on text or images, they train it on the process of optimization itself. It learns the *strategy* of searching for a good answer.
Jane: You got it. And the implications are pretty big. Think about all the engineering time that goes into hand-crafting optimizers for every new task. If this works, you could have one general-purpose tool that just works out of the box.
Tom: And that’s what we’re going to dig into. How did they build this thing, and does it actually hold up against the old guard like CMA-ES? Stick around.
Paper discussion summary: Tom: Alright, so we’ve set the stage. The paper is “Pretrained Optimization Model for Zero-Shot Black Box Optimization,” and Jane, you mentioned it learns a strategy. But what does that actually mean in practice?
Jane: So imagine you have a population of candidate solutions, like a swarm of points in space. The model looks at all of them, sees which ones are doing well and which are failing, and then decides how to mix them together to create the next generation.
Tom: So it’s like an evolutionary process, but the rules of evolution are learned, not hardcoded.
Jane: Precisely. The authors designed two main modules. One is a “Learned Mutation Module,” which decides how to blend individuals together to create new candidates. The other is a “Learned Crossover Module,” which decides how much of a new candidate should replace the old one.
Tom: And the clever part is that these modules are trainable. They used a method called MetaGBT to train the whole thing end-to-end on a set of simple, synthetic functions.
Jane: Right. They trained it on things like absolute value functions and Rosenbrock, just to learn the *mechanics* of optimization. Then they tested it on the BBOB benchmark, which is a suite of twenty-four very different, very tricky problems.
Tom: And the results are pretty wild. On the BBOB benchmark with thirty and one hundred dimensions, POM beat CMA-ES, which is considered the gold standard for this kind of problem. And the advantage got bigger as the dimensions went up.
Jane: That’s the part that got me. They trained it on ten-dimensional problems, and it works better than the classics on one hundred-dimensional problems. That’s not just generalization; that’s a whole new level of transfer learning.
Tom: They even tested it on robot control tasks, like teaching a bipedal walker to walk and an Enduro car to drive. POM was competitive or better than everything else there too.
Jane: And the coolest part for me is the few-shot ability. If you give POM just twenty-five random evaluations of the target problem to fine-tune it, you get a thirty percent performance improvement. That’s a tiny amount of data for a huge gain.
Tom: So it’s not just a zero-shot tool; it’s also a fantastic few-shot tool. That’s a double win.
Jane: Exactly. And that fine-tuning capability makes it practical for real-world use, where you might have a little bit of budget to spare.
Tom: So the summary is: a pretrained optimizer that learns the search strategy, beats the classics on hard benchmarks, and gets even better with a tiny bit of fine-tuning. What’s not to love?
Paper discussion improvements: Tom: We’re back with “Pretrained Optimization Model for Zero-Shot Black Box Optimization,” and Jane, we’ve covered the big wins. But what about the improvements the paper suggests? What did they actually change compared to previous attempts?
Jane: Great question. There have been other learned optimizers, like LES and LGA. But the authors point out that those often struggle with zero-shot performance. They’re weaker than CMA-ES when you just drop them on a new problem.
Tom: So what makes POM different? Why does it actually work?
Jane: The key improvement is the architecture and the training method. Previous methods either suffered from the “curse of dimensionality” in their model parameters, or they used reinforcement learning, which is notoriously unstable to train.
Tom: And POM sidesteps both of those issues.
Jane: Exactly. POM uses a gradient-based end-to-end training method called MetaGBT. That means the whole model is trained to directly minimize a loss function, which is a combination of making the population converge and keeping it diverse.
Tom: So it’s trained like a neural network, not like a reinforcement learning agent. That’s a much more stable and efficient way to learn.
Jane: And they also introduced a mask operation. During training, they randomly zero out parts of the mutation matrix. That forces the model to learn robust strategies that don’t rely on every individual talking to every other individual.
Tom: That’s a clever trick. It’s like teaching a team to work together even when some members are temporarily out of commission.
Jane: And it pays off. In their ablation study, when they removed the mask, the performance dropped significantly. It’s a crucial part of the design.
Tom: So the improvements are: a stable training method, a clever architecture, and a regularization trick that makes the whole thing more robust.
Jane: And the result is a model that not only beats the state-of-the-art but also scales well. They tested it on five hundred dimensions, and it still held its own.
Tom: So it’s not just a toy. It’s a serious tool that can handle high-dimensional, real-world problems.
Jane: And that’s what makes this paper so exciting. It’s not just an incremental improvement; it’s a fundamental shift in how we think about building optimizers.
Tom: So, what does the future hold? What’s the next step for this kind of research?
Jane: The authors mention that the relationship between model size and performance isn’t linear, and that larger models are harder to train. So there’s a lot of open questions about scaling.
Tom: And the time complexity of the attention mechanism is quadratic, which could be a bottleneck for very large populations. So there’s room for improvement there too.
Jane: But even with those limitations, this is a huge step forward. It shows that we can learn optimization itself, and that’s a powerful idea.
Conclusion: Tom: Alright, we’ve reached the end of our time with “Pretrained Optimization Model for Zero-Shot Black Box Optimization.” Jane, give us the final takeaway.
Jane: The takeaway is that we now have a pretrained optimizer that can be applied to new problems without any tuning, and it beats the classic algorithms. It learns the search strategy from simple training tasks and then transfers that knowledge to complex, high-dimensional problems.
Tom: And it’s not just about beating benchmarks. It’s about changing the workflow. Instead of hand-crafting an optimizer for every new task, you just load POM and go.
Jane: And if you have a little bit of budget, you can fine-tune it and get even better results. That’s a practical tool that engineers can actually use.
Tom: So, what’s the big picture here? Where does this leave the field?
Jane: It leaves us with a new paradigm. We’re moving from designing algorithms to learning them. And that’s a shift that could have a huge impact on everything from hyperparameter tuning to robot control to scientific discovery.
Tom: And it’s a reminder that the tools we use to solve problems can themselves be optimized.
Jane: Well said, Tom. This paper is a great example of that idea in action. We’ll be watching to see where this line of research goes next.
Tom: Thanks for joining us, everyone. We’ll see you next time with another paper from the arXiv.
Jane: Take care, and keep optimizing.
Xiaobin Li, Yujian Betterrest Li, Kai Wu, Xiaoyu Zhang, Handing Wang, Jing Liu
Xidian University
cs.NE, cs.AI
Submitted: 2026-08-11
Updated: 2026-08-12
Code: https://github.com/ninja-wm/POM
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 57/100
The gist: The paper introduces a Pretrained Optimization Model (POM) designed to address the challenge of zero-shot black-box optimization (BBO).
Terminology
Summary
The paper introduces a Pretrained Optimization Model (POM) designed to address the challenge of zero-shot black-box optimization (BBO). The authors define zero-shot optimization as optimizing a target task that was not seen during training, aiming to provide the optimal solution without or with minimal adjustments to the optimizer.
They note that current optimizers often struggle with zero-shot optimization and require intricate hyperparameter tuning to adapt to new tasks.
The paper's contributions are summarized as follows:
-
Excellent ability to solve zero-shot BBO
: POM demonstratesa substantial performance advantage over state-of-the-art black-box optimizers.
-
Excellent ability to solve few-shot BBO
: With25 random function evaluations,
POM achievesmore than 30% performance improvement
through fine-tuning.
The paper formalizes a black-box optimization problem as a minimization problem: min f(x), s.t. xi ∈ [li, ui], where x = (x1, x2,..., xd) represents the solution, l and u are lower and upper bounds, and d is the dimension. Two key definitions are provided:
-
Zero-shot Optimization: "an optimizer that is applied directly to solve a continuous black-box optimization problem f without any tuning... does not require any contextual information about f and can be directly used to handle problems of any dimensionality."
-
Few-shot Optimization:
it is permissible to fine-tune the optimizer using a small portion of the function evaluation budget for the objective task, and then use the fine-tuned optimizer to solve f.
POM consists of three main components:
LMM generates candidate solutions through the equation Vt = St × Xt, where Xt is the population at generation t and St ∈ RN×N is a mutation strategy matrix that evolves across generations. The module is designed based on Multi-head self-attention (MSA)
and uses population information including normalized fitness and centralized ranking of individuals. The computation involves:
-
Ĥt = Tanh(Ht × Wm1 + bm1)
-
Qt = Tanh(Ĥt × Wm2 + bm2)
-
Kt = Tanh(Ĥt × Wm3 + bm3)
-
Ŝt = Tanh(Qt × (Kt)T / √(dm))
A mask operation is applied where the probability of setting each element in Ŝt to 0 is rmask,
enhancing POM's ability to learn efficient and robust strategies.
LCM adaptively generates crossover probabilities crt based on population information. It is designed based on FFN
and uses the gumbel softmax method to handle the non-differentiable nature of discrete crossover operations. The crossover operation is performed as:
- uti = cvti,0 · xti + cvti,1 · vit
SM executes a 1-to-1 selection strategy
between Ut and Xt to produce the next-generation population Xt+1.
The paper introduces MetaGBT (Meta Gradient-Based Training), an end-to-end gradient-based training method
ensuring stable and rapid training for POM.
The training process involves:
-
Training on a set of functions TF1-TF5 with diverse landscape features including
Unimodal,
Separable,
Multimodal,
Non-separable,
Rotated,
andAsymmetrical
characteristics. -
Using a normalized loss function: li = li1 − λli2, where li1 encourages convergence and li2 encourages population diversity, with λ set to 0.005.
POM was evaluated on 24 BBOB functions with dimensions d = 30 and d = 100. The results show that POM significantly outperforms all methods, showcasing its efficacy across varying dimensions.
Despite being trained solely on TF1-TF4 with d = 10, POM excels in higher dimensions (d = 30, 100, 500), with its performance advantage becoming more pronounced with increasing dimensionality.
-
Bipedal Walker: POM
achieves stable and swift convergence, ultimately attaining the highest score
on a task with d = 874 parameters. -
Enduro: POM
maintains a superior balance between exploration and exploitation, outperforming LSHADE
on a task with d = 4149 parameters.
The ablation study shows that "the negative impact on POM's performance ranked as follows: NO MASK > NO LMM > UNTRAINED > NO LCM," demonstrating that all modules contribute to overall performance.
Fine-tuning POM with a small number of samples leads to significant performance improvements even with a small sample size.
The paper explores POM at different scales (VS, S, M, L, VL, XL) and finds that XL achieves the best performance,
with principles derived: "1) Larger models can have stronger capabilities but are more challenging to train; 2) Training difficulty and model scale do not exhibit a simple linear relationship; 3) Larger models require more functions for effective training."
Key observations include: "1) Generally, superior individuals receive higher weights during LMM... 2) Across diverse function problems, POM dynamically generates optimization strategies... 3) Disadvantaged individuals exhibit a more uniform weight distribution, potentially aiding in their escape from local optima."
LCM displays the capacity to adaptively generate diverse strategies for individuals across different ranks in the population,
with top-ranking individuals within the top 20... exhibiting a flexible crossover strategy
and lower-ranking individuals showing an increasing overall probability of crossover, promoting exploration.
The paper acknowledges limitations including: Model size: the relationship between the model size and the performance of POM is not a strict linear relationship
and Time performance: We introduced an operation similar to the attention mechanism, whose time complexity is O(n2), which makes POM require a lot of time cost when processing large-scale populations.
The paper concludes that POM, a novel Pretrained Optimization Model designed to address the inefficiencies of existing methods in zero-shot optimization,
demonstrates superiority over other black-box optimizers, particularly in high-dimensional scenarios
and excels in solving few-shot optimization problems.
Future research directions include designing enhanced loss functions to optimize POM for both population convergence and diversity
and addressing the limitations of model scale and time performance.
Improvements for AI systems
Based on the scientific paper, here are the specific improvements I can make to AI systems and the resulting capabilities:
Improvement: Integrate POM as a pre-trained optimization layer that can be directly applied to unseen optimization tasks without hyperparameter tuning.
Capabilities:
-
Optimize any continuous black-box function (e.g., hyperparameter tuning, neural architecture search) without prior knowledge of the task
-
Handle problems of varying dimensionality (10 to 500+ dimensions) despite being trained on only 10-dimensional tasks
-
Outperform CMA-ES, L-SHADE, and other state-of-the-art optimizers on high-dimensional problems (30–500 dimensions)
Improvement: Implement POM's fine-tuning capability that uses a small budget of function evaluations to adapt the optimizer to a specific target task.
Improvement: Replace fixed mutation/crossover operators in evolutionary algorithms with POM's learnable modules (LMM for mutation, LCM for crossover) that dynamically generate strategies based on population state.
Improvement: Use MetaGBT (Meta Gradient-Based Training) to train optimization models across diverse task distributions in a stable, end-to-end manner.
Improvement: Deploy POM as a drop-in replacement for traditional evolutionary algorithms in reinforcement learning and control tasks.
Improvement: Use POM's interpretable strategy matrices (S t for mutation, cr t for crossover) to analyze and debug optimization behavior.
Improvement: Leverage POM's ability to generalize to dimensions far beyond its training dimension (10 → 500) for real-world high-dimensional problems.
Improvement: Use POM's scale-performance analysis to guide model architecture selection for different optimization budgets.
The improved AI system can:
-
Optimize any continuous black-box function without prior tuning, across dimensions from 10 to 500+
-
Adapt to new tasks with minimal samples (few-shot) for rapid deployment
-
Automatically design mutation and crossover strategies based on population dynamics
-
Train stable, generalizable optimizers via MetaGBT without reinforcement learning instability
-
Outperform state-of-the-art optimizers (CMA-ES, L-SHADE, TurBO) on high-dimensional and real-world tasks
-
Provide interpretable optimization strategies for debugging and analysis
-
Scale efficiently to large populations and high-dimensional problems with minimal time cost
Sources
- Deep Neuroevolution: Genetic Algorithms Are a Competitive Alternative for Training Deep Neural Networks for Reinforcement Learning
- Generative Pretraining for Black-Box Optimization
- Diffusion Models for Black-Box Optimization
- The CMA Evolution Strategy: A Tutorial
- Meta Learning Black-Box Population-Based Optimizers
- On the Relationship Between the OpenAI Evolution Strategy and Stochastic Gradient Descent
- Language Model Crossover: Variation through Few-Shot Prompting
- Algorithm Evolution Using Large Language Model
- Large Language Models as Optimizers
- Eureka: Human-Level Reward Design via Coding Large Language Models
- EvoPrompting: Language Models for Code-Level Neural Architecture Search
- LLMatic: Neural Architecture Search via Large Language Models and Quality Diversity Optimization
- Exploring the True Potential: Evaluating the Black-box Optimization Capability of Large Language Models
- LLaMoCo: Instruction Tuning of Large Language Models for Optimization Code Generation
- Evolution of Heuristics: Towards Efficient Automatic Algorithm Design Using Large Language Model
- Transformer Neural Processes: Uncertainty-Aware Meta Learning Via Sequence Modeling
- LICO: Large Language Models for In-Context Molecular Optimization
- Categorical Reparameterization with Gumbel-Softmax
- Adam: A Method for Stochastic Optimization
- Comparison of High-Dimensional Bayesian Optimization Algorithms on BBOB
Related papers
- Evolutionary Ensemble of Agents
- Encoding and Decoding Temporal Signals with Spiking Bandpass Wavelets
- Large Language Models and Evolutionary Computation: A Critical Review of Bidirectional Interaction, Automated Algorithm Design, and Co-Adaptive Systems
- Learning Alzheimer's Disease Signatures by bridging EEG with Spiking Neural Networks and Biophysical Simulations
- Investigating Hyperparameter Optimization and Transferability for ES-HyperNEAT: A TPE Approach
- S-AI-Recursive: A Bio-Inspired and Temporal Sparse AI Architecture for Iterative, Introspective, and Energy-Frugal Reasoning