A Logarithmic Regret Bound for Optimistic Hedge in General-Sum Games
Junsoo Ha
Abstract
Can simple no-regret dynamics attain smaller regret in self-play than against arbitrary adversaries? In n-player general-sum games, Daskalakis et al. 2021 proved an O(n di4 T) individual regret bound for Optimistic Hedge, which improves upon the classical O( T) adversarial regret bound. In this work, we show that Optimistic Hedge with a constant step size can further achieve O( n di T) individual external regret under expected loss-vector feedback. The time-averaged play consequently enjoys a coarse correlated equilibrium gap O( n d T/T), where d=i di. The improvement comes from a larger admissible step size η=Θ(1/( n T)). Our analysis proves factorial bounds on high-order differences of probability-weighted pairwise loss gaps, then applies finite-difference interpolation in a fixed Euclidean norm. These estimates sharpen the analysis of Daskalakis et al. 2021 and yield a logarithmic regret bound.
Create a lesson
Related papers
A Nearly Tight Lower Bound for Matroid Intersection Prophet Inequalities
Dimitris Fotakis, Charalampos Platanos, Thanos Tolias
Faster Verification of PJR+ via Mincuts
Drew Springham
Exact Regret Frontiers and Externality Scheduling in Centralized Serial-Dictatorship Bandits
Lishang Xu, Guodong Ma, Pengcheng Weng et al.
Condorcet-type properties of the linear ordering problem with ties
Daichi Kawashima, Noriyoshi Sukegawa
On Periodic and Aperiodic Optimal Strategies in Solvency Games
Quentin Guilmant, Florian Luca, Richard Mayr et al.
Efficient Nash Equilibrium Computation for Cybersecurity Games
Michael Lanier, David Farmer, Yevgeniy Vorobeychik