Exact Regret Frontiers and Externality Scheduling in Centralized Serial-Dictatorship Bandits
Lishang Xu, Guodong Ma, Pengcheng Weng, Zixuan Xia
Abstract
Exploration in centralized serial-dictatorship matching bandits must use complete matchings, so learning one player--arm pair can impose regret on others. We study this externality under a known common priority order and Gaussian rewards with unit variance. We show that the matching-level Graves--Lai constraints reduce to finitely many pairwise exploration quotas and, at top-choice-separated instances, yield a polynomial-size marginal linear program. At these instances, the exact attainable set of expected logarithmic regret coefficients is G(θ)(θ), where is the feasible matching-allocation set and G maps allocations to player regret. The usual upper-closed Graves--Lai region can be strictly larger despite having the same Pareto-minimal boundary. We further show that identical exploration quotas can induce very different regret through their scheduling. Finally, we construct estimate--solve--track policies, uniformly good on the full row-strict class, that attain every fixed positively weighted optimum without assuming optimizer uniqueness. Every Pareto-minimal point is pointwise attainable, possibly through an instance-calibrated target.
Create a lesson
Related papers
A Nearly Tight Lower Bound for Matroid Intersection Prophet Inequalities
Dimitris Fotakis, Charalampos Platanos, Thanos Tolias
Faster Verification of PJR+ via Mincuts
Drew Springham
A Logarithmic Regret Bound for Optimistic Hedge in General-Sum Games
Junsoo Ha
Condorcet-type properties of the linear ordering problem with ties
Daichi Kawashima, Noriyoshi Sukegawa
On Periodic and Aperiodic Optimal Strategies in Solvency Games
Quentin Guilmant, Florian Luca, Richard Mayr et al.
Efficient Nash Equilibrium Computation for Cybersecurity Games
Michael Lanier, David Farmer, Yevgeniy Vorobeychik