Skip to content

Toward Optimal Second-Order Path-Length Guarantee for Adversarial Multi-Armed Bandits

Mengxiao Zhang

cs.LGarXiv:2608.15996

Abstract

We study second-order path-length regret in adversarial K-armed bandits against oblivious loss sequences. Bubeck et al. [2019] designed an algorithm that achieves O(K+KQ∞,1) regret, where Q∞,1 is the first-order path length, and left open whether O(poly(K)1+Q∞,2) regret is achievable under bandit feedback, where Q∞,2 is the second-order path length. Somewhat surprisingly, we resolve this question positively by showing that with a more involved analysis, the exact same algorithm of Bubeck et al. [2019] achieves O(K(KT)+K(KT)(1+Q∞,2)) expected regret when Q∞,2 is known, where T is the horizon. This matches the Ω(KQ∞,2) lower bound up to logarithmic factors and additive terms. We further remove the knowledge of Q∞,2 using an adaptive restart scheme whose path-length estimator has uniformly bounded increments.

Create a lesson