Sublogarithmic Swap Regret in Multiplayer General-Sum Games via Hybrid Regularization
Taira Tsuchiya
Abstract
Swap regret governs the rate at which uncoupled learning dynamics converge to correlated equilibria in multiplayer general-sum games. Under full-information feedback, the best previous guarantee when every player follows the same dynamics grows logarithmically in the horizon T. We construct uncoupled dynamics under which every player incurs only O(nm2 m T) swap regret, where n is the number of players and m bounds the number of actions per player. To our knowledge, this is the first sublogarithmic individual guarantee in this setting, and it implies that the time-averaged product distribution of play is an O(nm2 m T/T)-approximate correlated equilibrium. The key algorithmic choice is to combine the Blum--Mansour reduction with optimistic follow-the-regularized-leader using a hybrid regularizer that separately weights negative Shannon entropy and the log-barrier: the entropy controls the optimistic prediction error, whereas the log-barrier controls the transition-matrix movement through its Bregman divergence. A new sensitivity theorem for stationary distributions of Markov chains, which involves neither mixing parameters nor the smallest transition probability, transfers this control to the played strategies and yields a simpler analysis without local-norm or self-concordance arguments. The guarantee is preserved by an adversarially robust variant that additionally ensures O(nm2 m T+mT m) swap regret against arbitrary utility sequences, and by a horizon-free variant that requires no prior knowledge of T.
Create a lesson
Related papers
On the Role of Tie-Breaking Rules in the Convergence of Fictitious Play for Symmetric First-Price Auctions
Benjamin Heymann
Epsilon-Nash Equilibria in History-Dependent SA-MDPs
Brandon Gary Kaplowitz, Dominik Bohnet Zurcher, Akash Agrawal et al.
Core stability recognition for minimum-cost spanning tree games: Parameterized perspective
Michal Dvořák, Ioannis Kakatelis, Dušan Knop
Second-Best Gains from Trade in Matching Markets
Xiaohui Bei, Bo Li, Wenhao Wu et al.
Equilibria of Round-Robin: Computational Hardness and Fairness for Few Subadditive Agents
Paul W. Goldberg, Alexandros Hollender, Giannis Tyrovolas
Estimate then Predict: Convex Formulation for Travel Demand Forecasting
Youngseo Kim, Gioele Zardini, Samitha Samaranayake et al.