Toward Optimal Second-Order Path-Length Guarantee for Adversarial Multi-Armed Bandits
Mengxiao Zhang
Abstract
We study second-order path-length regret in adversarial K-armed bandits against oblivious loss sequences. Bubeck et al. [2019] designed an algorithm that achieves O(K+KQ∞,1) regret, where Q∞,1 is the first-order path length, and left open whether O(poly(K)1+Q∞,2) regret is achievable under bandit feedback, where Q∞,2 is the second-order path length. Somewhat surprisingly, we resolve this question positively by showing that with a more involved analysis, the exact same algorithm of Bubeck et al. [2019] achieves O(K(KT)+K(KT)(1+Q∞,2)) expected regret when Q∞,2 is known, where T is the horizon. This matches the Ω(KQ∞,2) lower bound up to logarithmic factors and additive terms. We further remove the knowledge of Q∞,2 using an adaptive restart scheme whose path-length estimator has uniformly bounded increments.
Create a lesson
Related papers
How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents
Zixi Chen, Akshay Vegesna, Samip Dahal et al.
Evidence-Grounded Agentic Formulation Development in an Autonomous Laboratory
Michael M. Craig, Riley J. Hickman, Yingshan Ma et al.
Probabilistic Linear Explanations
Frederic Koriche, Jean-Marie Lagniez, Chi Tran
Double descent is the principle of least action
Congzhou M Sha
RLLBC-Lib: An Educational Code Library for Reinforcement Learning and Learning-Based Control
Bernd Frauenknecht, Emma Cramer, Artur Eisele et al.
Higher-order pruning of experts in mixture-of-experts language models
Alex M. Tseng, Prannay Kaul, Luca Zancato et al.