Skip to content

An Efficient Minimax-Optimal Algorithm for Adversarial m-Set Bandits

Francesco Bacchiocchi, Tommaso Cesari, Roberto Colomboni

cs.LGarXiv:2608.12231

Abstract

We study adversarial combinatorial bandits with m-set actions, where at each round the learner selects m out of d items and observes only the aggregate loss of the selected items. The resulting action set contains K=dm elements and can therefore be exponentially large. Nevertheless, the loss of every action is determined by the same d-dimensional vector of item losses. We propose a computationally efficient algorithm that exploits this structure without explicitly enumerating the action set. Against adaptive non-anticipating adversaries, it guarantees, with probability at least 1-δ, regret against the best fixed action of \[ RT = O(dT(K/δ)). \] This matches the high-probability regret bound of the finite-action EXP3-KW algorithm of Zimmert and Lattimore, whose direct implementation may require exponential space. Our algorithm instead represents each sampling distribution with d parameters and runs in polynomial time without enumerating the action set. Thus, it resolves the open problem posed by Maiti et al. We complement this upper bound with a matching high-probability lower bound. For all sufficiently small δ, every randomized policy admits a deterministic adaptive non-anticipating adversary for which, with probability at least δ, \[ RT = Ω(dT(K/δ)). \] Thus, the rate is minimax optimal up to universal constants in this regime. In particular, setting m=1 proves that the K for ordinary K-armed bandits against adaptive non-anticipating adversaries is unavoidable, closing the remaining K gap between confidence-tuned upper and lower bounds left by Gerchinovitz and Lattimore.

Create a lesson