Independent Reinforcement Learning in Discounted Markov Games
Asrin Efe Yorulmaz, Ugur Aydin, Tamer Basar
Abstract
In this work, we study radically uncoupled learning in discounted general-sum Markov games. Assuming ``ETH for PPAD", we show that, for every fixed discount factor, there is no polynomial-time algorithm for computing inverse-polynomially accurate coarse correlated equilibria in discounted general-sum Markov games when players learn independently in decentralized settings. Complementing this hardness result, we provide what appears to be the first radically uncoupled algorithm with sub-exponential convergence guarantees to coarse correlated equilibria in discounted general-sum Markov games without imposing any structural restrictions on the game. Our algorithm is a layered variant of optimistic mirror descent with an increasing step-size schedule tailored to the multi-agent setting. Finally, we develop both full-feedback and partial feedback versions of the aforementioned algorithm and establish sub-exponential convergence guarantees for each case.
Create a lesson
Related papers
Approximately Efficient Multidimensional Bilateral Trade
Aviad Rubinstein, Xizhi Tan, Zixin Zhou
Almost Envy-Freeness for Additive Mixed Manna with Entitlements: Deterministic and Randomized Guarantees
Zehan Lin, Shengxin Liu, Biaoshuai Tao et al.
Fair Stable Matching: A Nash Social Welfare Approach
Parth Desai, Rasheed M, Ganesh Ghalme et al.
The Complexity of Justified Representation with Additive Utilities
Carmel Baharav, Jakob de Raaij, Agnès Totschnig
Weighted Fair Division of Indivisible Mixed Manna
Nicholas Teh
Graph Coloring with Color Preferences
Tomohiro Koana, Yeeseok Oh, Hirotaka Yoneda