Skip to content

A Logarithmic Regret Bound for Optimistic Hedge in General-Sum Games

Junsoo Ha

cs.GTarXiv:2609.19677

Abstract

Can simple no-regret dynamics attain smaller regret in self-play than against arbitrary adversaries? In n-player general-sum games, Daskalakis et al. 2021 proved an O(n di4 T) individual regret bound for Optimistic Hedge, which improves upon the classical O( T) adversarial regret bound. In this work, we show that Optimistic Hedge with a constant step size can further achieve O( n di T) individual external regret under expected loss-vector feedback. The time-averaged play consequently enjoys a coarse correlated equilibrium gap O( n d T/T), where d=i di. The improvement comes from a larger admissible step size η=Θ(1/( n T)). Our analysis proves factorial bounds on high-order differences of probability-weighted pairwise loss gaps, then applies finite-difference interpolation in a fixed Euclidean norm. These estimates sharpen the analysis of Daskalakis et al. 2021 and yield a logarithmic regret bound.

Create a lesson