Fetching the paper…
Reading the bibliography…
We study infinite-horizon discounted two-player zero-sum Markov games, and develop a decentralized algorithm that provably converges to the set of Nash equilibria under self-play.
Stochastic games
Lloyd S Shapley · 1953
Earlier work this paper cites.
On nonterminating stochastic games
Alan J Hoffman and Richard M Karp · 1966
Earlier work this paper cites.
Algorithms for stochastic games with geometrical interpretation
MA Pollatschek and B Avi-Itzhak · 1969
Earlier work this paper cites.
Discounted markov games: Generalized policy iteration method
J Van Der Wal · 1978
Earlier work this paper cites.
A modification of the arrow-hurwicz method for search of saddle points
Leonid Denisovich Popov · 1980
Earlier work this paper cites.
On the algorithm of pollatschek and avi-ltzhak
Jerzy A Filar and Boleslaw Tolwinski · 1991
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
A unified analysis of value-function-based reinforcement-learning algorithms
Csaba Szepesvári and Michael L Littman · 1999
Earlier work this paper cites.
Rational and convergent learning in stochastic games
Michael Bowling and Manuela Veloso · 2001
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
Peter Auer and Ronald Ortner · 2007
Earlier work this paper cites.
Awesome: A general multiagent learning algorithm that converges in self-play and learns a best response against stationary opponents
Vincent Conitzer and Tuomas Sandholm · 2007
Earlier work this paper cites.
Online optimization with gradual variations
Chao-Kai Chiang, Tianbao Yang, Chia-Jung Lee, Mehrdad Mahdavi, Chi-Jen Lu, Rong Jin, and Shenghuo Zhu · 2012
Earlier work this paper cites.
Competitive Markov decision processes
Jerzy Filar and Koos Vrieze · 2012
Earlier work this paper cites.
First-order algorithm with 𝒪 ( ln ( 1 / ϵ ) ) \mathcal{O}(\ln(1/\epsilon)) convergence for ϵ \epsilon -equilibrium in two-person zero-sum games
Andrew Gilpin, Javier Pena, and Tuomas Sandholm · 2012
Earlier work this paper cites.
Optimization, learning, and games with predictable sequences
Sasha Rakhlin and Karthik Sridharan · 2013
Cited alongside, same era.
Approximate dynamic programming for two-player zero-sum markov games
Julien Perolat, Bruno Scherrer, Bilal Piot, and Olivier Pietquin · 2015
Cited alongside, same era.
Fast convergence of regularized learning in games
Vasilis Syrgkanis, Alekh Agarwal, Haipeng Luo, and Robert E Schapire · 2015
Cited alongside, same era.
Softened approximate policy iteration for markov games
Julien Pérolat, Bilal Piot, Matthieu Geist, Bruno Scherrer, and Olivier Pietquin · 2016
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
Online reinforcement learning in stochastic games
Chen-Yu Wei, Yi-Te Hong, and Chi-Jen Lu · 2017
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2020
Later among the works it cites.
Provable self-play algorithms for competitive reinforcement learning
Yu Bai and Chi Jin · 2020
Later among the works it cites.
Near-optimal reinforcement learning with self-play
Yu Bai, Chi Jin, and Tiancheng Yu · 2020
Later among the works it cites.
Independent policy gradient methods for competitive reinforcement learning
Constantinos Daskalakis, Dylan J. Foster, and Noah Golowich · 2020
Later among the works it cites.
Tight last-iterate convergence rates for no-regret learning in multi-player games
Noah Golowich, Sarath Pattathil, and Constantinos Daskalakis · 2020
Later among the works it cites.
Linear last-iterate convergence for matrix games and stochastic games
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Cited alongside, same era.
Actor-critic fictitious play in simultaneous move multistage games
Julien Perolat, Bilal Piot, and Olivier Pietquin · 2018
Cited alongside, same era.
Politex: Regret bounds for policy iteration using expert prediction
Yasin Abbasi-Yadkori, Peter Bartlett, Kush Bhatia, Nevena Lazic, Csaba Szepesvari, and Gellért Weisz · 2019
Cited alongside, same era.
On the convergence of single-call stochastic extra-gradient methods
Yu-Guan Hsieh, Franck Iutzeler, Jérôme Malick, and Panayotis Mertikopoulos · 2019
Cited alongside, same era.
Interaction matters: A note on non-asymptotic local convergence of generative adversarial networks
Tengyuan Liang and James Stokes · 2019
Cited alongside, same era.
Learning to collaborate in markov decision processes
Goran Radanovic, Rati Devidze, David Parkes, and Adish Singla · 2019
Cited alongside, same era.
Chung-Wei Lee, Haipeng Luo, Chen-Yu Wei, and Mengxiao Zhang · 2020
Later among the works it cites.
Policy optimization in zero-sum markov games: Fictitious self-play provably attains nash equilibria, 2020
Boyi Liu, Zhuoran Yang, and Zhaoran Wang · 2020
Later among the works it cites.
A unified analysis of extra-gradient and optimistic gradient methods for saddle point problems: Proximal point approach
Aryan Mokhtari, Asuman Ozdaglar, and Sarath Pattathil · 2020
Later among the works it cites.
Fictitious play in zero-sum stochastic games
Muhammed O Sayin, Francesca Parise, and Asuman Ozdaglar · 2020
Later among the works it cites.
Solving discounted stochastic two-player games with near-optimal time and sample complexity
Aaron Sidford, Mengdi Wang, Lin Yang, and Yinyu Ye · 2020
Later among the works it cites.
Learning zero-sum simultaneous-move markov games using function approximation and correlated equilibrium
Qiaomin Xie, Yudong Chen, Zhaoran Wang, and Zhuoran Yang · 2020
Later among the works it cites.
A sharp analysis of model-based reinforcement learning with self-play
Qinghua Liu, Tiancheng Yu, Yu Bai, and Chi Jin · 2021
Closest in time.
Online learning in unknown markov games
Yi Tian, Yuanhao Wang, Tiancheng Yu, and Suvrit Sra · 2021
Closest in time.
Linear last-iterate convergence in constrained saddle-point optimization
Chen-Yu Wei, Chung-Wei Lee, Mengxiao Zhang, and Haipeng Luo · 2021
Closest in time.