Reinforcement learning to choose optimizers
Martin van der Schelling, Deepesh Toshniwal, Miguel A. Bessa
Abstract
No single optimization method is uniformly best for all problems, and the most suitable optimizer choice can change during a run. Existing approaches that change optimizer during execution typically predetermine part of the strategy: the portfolio is restricted to one algorithm class, the switch occurs once at a fixed time, or the frequency of decisions is treated as a hyperparameter rather than a learned one. We introduce "Reinforcement Learning to Choose Optimizers", which formulates the optimization algorithm choice as a sequential decision-making problem. At each decision, a recurrent policy reads the current run state and decides both which optimizer should be used next and for how long. The portfolio includes both gradient-based and derivative-free optimizers, and each switch passes on the current best solution and a representative step size. A context proxy conditions a gating network over expert heads, and training employs a decoupled actor-critic whose return is expressed in the same empirical runtime distribution metric used at evaluation. Training tasks and portfolio are designed jointly so that no optimizer dominates. On unseen problems, the learned policy outperforms every portfolio optimizer at all but the smallest budgets, and it remains robust under distribution shift.
Create a lesson
Related papers
Model-Free Surrogate-Assisted Neural Architecture Search for Evolving Variable-Length Dense Blocks
Asif Ameer, Maryam Bashir, Irfan Younas et al.
Semantics-Guided Automatic Tensorization for Multiobjective Evolutionary Algorithms: A Multi-Agent Framework
Zhenyu Liang, Beichen Huang, Bowen Zheng et al.
LLM-Driven Joint Evolution of Coupled Heuristics Components for Routing Optimization
Juntao Wei, Yangming Zhou, Zhibin Jiang et al.
Memory as an Energy Landscape---Hopfield
Nima Dehghani
CircuitsDNA: Discovering Unconventional Multi-Accuracy Arithmetic Circuits via Evolutionary Synthesis
Ruichen Qi, Junyi Luo, Xinting Jiang et al.
Investigating Hyperparameter Optimization and Transferability for ES-HyperNEAT: A TPE Approach
Romain Claret, Michael O'Neill, Paul Cotofrei et al.