Reinforcement learning to choose optimizers
cs.NE, cs.LG, math.OC
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: 20 pages, 14 figures
Code: https://github.com/bessagroup/rl2co
License: http://creativecommons.org/licenses/by/4.0/
The gist: No single optimization method is uniformly best for all problems, and the most suitable optimizer choice can change during a run.
Terminology
Abstract
No single optimization method is uniformly best for all problems, and the most suitable optimizer choice can change during a run. Existing approaches that change optimizer during execution typically predetermine part of the strategy: the portfolio is restricted to one algorithm class, the switch occurs once at a fixed time, or the frequency of decisions is treated as a hyperparameter rather than a learned one. We introduce "Reinforcement Learning to Choose Optimizers", which formulates the optimization algorithm choice as a sequential decision-making problem. At each decision, a recurrent policy reads the current run state and decides both which optimizer should be used next and for how long. The portfolio includes both gradient-based and derivative-free optimizers, and each switch passes on the current best solution and a representative step size. A context proxy conditions a gating network over expert heads, and training employs a decoupled actor-critic whose return is expressed in the same empirical runtime distribution metric used at evaluation. Training tasks and portfolio are designed jointly so that no optimizer dominates. On unseen problems, the learned policy outperforms every portfolio optimizer at all but the smallest budgets, and it remains robust under distribution shift.
Sources
- Automated Algorithm Selection: Survey and Perspectives
- Learning to learn by gradient descent by gradient descent
- Learning to Optimize
- Switching between Numerical Black-box Optimization Algorithms with Warm-starting Policies
- Deep Reinforcement Learning for Dynamic Algorithm Selection: A Proof-of-Principle Study on Differential Evolution
- R2 Indicator and Deep Reinforcement Learning Enhanced Adaptive Multi-Objective Evolutionary Algorithm
- Black-Box Optimization Revisited: Improving Algorithm Selection Wizards through Massive Benchmarking
- MetaBox-v2: A Unified Benchmark Platform for Meta-Black-Box Optimization
- Toward Automated Algorithm Design: A Survey and Practical Guide to Meta-Black-Box-Optimization
- Exploratory Landscape Analysis is Strongly Sensitive to the Sampling Strategy
- On the Utility of Probing Trajectories for Algorithm-Selection
- DoE2Vec: Deep-learning Based Features for Exploratory Landscape Analysis
- When Switching Algorithms Helps: A Theoretical Study of Online Algorithm Selection
- Understanding and correcting pathologies in the training of learned optimizers
- A Closer Look at Learned Optimization: Stability, Robustness, and Inductive Biases
- To Switch or not to Switch: Predicting the Benefit of Switching between Algorithms based on Trajectory Features
- Per-run Algorithm Selection with Warm-starting using Trajectory-based Features
- Trajectory-based Algorithm Selection with Warm-starting
- Improving Generalization Performance by Switching from Adam to SGD
- Task-free Adaptive Meta Black-box Optimization
Related papers
- Evolutionary Ensemble of Agents
- Encoding and Decoding Temporal Signals with Spiking Bandpass Wavelets
- Large Language Models and Evolutionary Computation: A Critical Review of Bidirectional Interaction, Automated Algorithm Design, and Co-Adaptive Systems
- Learning Alzheimer's Disease Signatures by bridging EEG with Spiking Neural Networks and Biophysical Simulations
- Investigating Hyperparameter Optimization and Transferability for ES-HyperNEAT: A TPE Approach
- S-AI-Recursive: A Bio-Inspired and Temporal Sparse AI Architecture for Iterative, Introspective, and Energy-Frugal Reasoning