Skip to content

Toward the Optimal Regret-Instability Trade-off in Multi-Armed Bandits

Kaifei Wang, Yinyu Ye, Han Zhong

stat.MLarXiv:2608.17841

Abstract

Multi-armed bandit algorithms are evaluated by regret, yet comparable regret can coexist with different allocations across independent runs. We study the trade-off between worst-case regret RK,T and instability SK,T, defined as the largest standard deviation of a terminal pull count, for K arms and T rounds. We prove the finite-time lower bound RK,T SK,T C T3/2, where C is independent of K and T, under a finite-time regret condition and without the regularity assumptions imposed in the prior asymptotic analysis. We also introduce Stabilized Lower-Envelope UCB (SLE-UCB), a new tunable algorithm combining a running lower-envelope index with a decreasing pull-count stabilizer. SLE-UCB satisfies RK,T SK,T=O(T3/2 K), with an implicit constant independent of K and T, matching the lower bound exactly in T and within a logarithmic factor in K. To prove the instability bound, we develop a new offline top-prefix representation that removes path dependence from online decisions. Together with single-reward perturbations and the Efron--Stein inequality, this representation controls pull-count variance. Thus, regret and instability depend reciprocally on K, while their product has no polynomial dependence on K. These results resolve the open question raised in the literature concerning the sharp arm-dependent regret--instability frontier.

Create a lesson