Fetching the paper…
Reading the bibliography…
This paper addresses the problem of learning an equilibrium efficiently in general-sum Markov games through decentralized multi-agent reinforcement learning.
Stochastic games
L. S. Shapley · 1953
Earlier work this paper cites.
Strategically zero-sum games: The class of games whose completely mixed equilibria cannot be improved upon
H. Moulin and J.-P. Vial · 1978
Earlier work this paper cites.
Team decision theory and information structures
Y.-C. Ho · 1980
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
A. S. Nemirovskij and D. B. Yudin · 1983
Earlier work this paper cites.
Correlated equilibrium as an expression of Bayesian rationality
R. J. Aumann · 1987
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
M. L. Littman · 1994
Earlier work this paper cites.
Gambling in a rigged casino: The adversarial multi-armed bandit problem
P. Auer, N. Cesa-Bianchi, Y. Freund, and R. E. Schapire · 1995
Earlier work this paper cites.
A generalized reinforcement-learning model: Convergence and applications
M. L. Littman and C. Szepesvári · 1996
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
C. Claus and C. Boutilier · 1998
Earlier work this paper cites.
A simple adaptive procedure leading to correlated equilibrium
S. Hart and A. Mas-Colell · 2000
Earlier work this paper cites.
Friend-or-Foe Q-learning in general-sum games
M. L. Littman · 2001
Earlier work this paper cites.
The complexity of decentralized control of Markov decision processes
D. S. Bernstein, R. Givan, N. Immerman, and S. Zilberstein · 2002
Earlier work this paper cites.
Reinforcement learning to play an optimal Nash equilibrium in team Markov games
X. Wang and T. Sandholm · 2002
Earlier work this paper cites.
Correlated-Q learning
A. Greenwald and K. Hall · 2003
Earlier work this paper cites.
Nash Q-learning for general-sum stochastic games
J. Hu and M. P. Wellman · 2003
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
M. Zinkevich · 2003
Earlier work this paper cites.
Individual Q-learning in normal form games
D. S. Leslie and E. J. Collins · 2005
Earlier work this paper cites.
Prediction, learning, and games
N. Cesa-Bianchi and G. Lugosi · 2006
Cited alongside, same era.
Computing correlated equilibria in multi-player games
C. H. Papadimitriou and T. Roughgarden · 2008
Cited alongside, same era.
Settling the complexity of computing two-player Nash equilibria
X. Chen, X. Deng, and S.-H. Teng · 2009
Cited alongside, same era.
The complexity of computing a Nash equilibrium
C. Daskalakis, P. W. Goldberg, and C. H. Papadimitriou · 2009
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
T. Jaksch, R. Ortner, and P. Auer · 2010
Cited alongside, same era.
Competitive Markov decision processes
J. Filar and K. Vrieze · 2012
Cited alongside, same era.
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
N. Brown and T. Sandholm · 2018
Later among the works it cites.
Is Q-learning provably efficient?
C. Jin, Z. Allen-Zhu, S. Bubeck, and M. I. Jordan · 2018
Later among the works it cites.
Scale-free online learning
F. Orabona and D. Pál · 2018
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al · 2019
Later among the works it cites.
Learning team-optimality for decentralized stochastic control and dynamic games
B. Yongacoglu, G. Arslan, and S. Yüksel · 2019
Later among the works it cites.
Provable self-play algorithms for competitive reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Kober, J. A. Bagnell, and J. Peters · 2013
Cited alongside, same era.
Stochastic networked control systems: Stabilization and optimization under information constraints
S. Yüksel and T. Başar · 2013
Cited alongside, same era.
Convex optimization: Algorithms and complexity
S. Bubeck et al · 2015
Cited alongside, same era.
Explore no more: Improved high-probability regret bounds for non-stochastic bandits
G. Neu · 2015
Cited alongside, same era.
Two-timescale algorithms for learning Nash equilibria in general-sum stochastic games
H. Prasad, P. LA, and S. Bhatnagar · 2015
Cited alongside, same era.
Decentralized Q-learning for stochastic teams and games
G. Arslan and S. Yüksel · 2016
Cited alongside, same era.
Y. Bai and C. Jin · 2020
Later among the works it cites.
Near-optimal reinforcement learning with self-play
Y. Bai, C. Jin, and T. Yu · 2020
Later among the works it cites.
Independent policy gradient methods for competitive reinforcement learning
C. Daskalakis, D. J. Foster, and N. Golowich · 2020
Later among the works it cites.
Online mirror descent and dual averaging: Keeping pace in the dynamic case
H. Fang, N. Harvey, V. Portella, and M. Friedlander · 2020
Later among the works it cites.
Solving discounted stochastic two-player games with near-optimal time and sample complexity
A. Sidford, M. Wang, L. Yang, and Y. Ye · 2020
Later among the works it cites.
Learning zero-sum simultaneous-move Markov games using function approximation and correlated equilibrium
Q. Xie, Y. Chen, Z. Wang, and Z. Yang · 2020
Later among the works it cites.
PAC reinforcement learning algorithm for general-sum Markov games
A. Zehfroosh and H. G. Tanner · 2020
Later among the works it cites.
A sharp analysis of model-based reinforcement learning with self-play
Q. Liu, T. Yu, Y. Bai, and C. Jin · 2021
Closest in time.
Online learning in unknown Markov games
Y. Tian, Y. Wang, T. Yu, and S. Sra · 2021
Closest in time.
Last-iterate convergence of decentralized optimistic gradient descent/ascent in infinite-horizon competitive Markov games
C.-Y. Wei, C.-W. Lee, M. Zhang, and H. Luo · 2021
Closest in time.
Provably efficient policy gradient methods for two-player zero-sum Markov games
Y. Zhao, Y. Tian, J. D. Lee, and S. S. Du · 2021
Closest in time.