Fetching the paper…
Reading the bibliography…
We obtain global, non-asymptotic convergence guarantees for independent learning algorithms in competitive reinforcement learning settings with two agents (i.e., zero-sum stochastic games).
Multi-agent reinforcement learning: A selective overview of theories and algorithms
Kaiqing Zhang, Zhuoran Yang, and Tamer Başar · 1911
Earlier work this paper cites.
A model of general economic equilibrium
John von Neumann · 1945
Earlier work this paper cites.
Stochastic Games
Lloyd Shapley · 1953
Earlier work this paper cites.
Convex analysis
R Tyrrell Rockafellar · 1970
Earlier work this paper cites.
The extragradient method for finding saddle points and other problems
GM Korpelevich · 1976
Earlier work this paper cites.
On algorithms for simple stochastic games
Anne Condon · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Ming Tan · 1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L. Littman · 1994
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Caroline Claus and Craig Boutilier · 1998
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
R-max–a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Nash Q-learning for general-sum stochastic games
Junling Hu and Michael P Wellman · 2003
Earlier work this paper cites.
Error bounds for approximate policy iteration
Rémi Munos · 2003
Earlier work this paper cites.
Finite-dimensional variational inequalities and complementarity problems
Francisco Facchinei and Jong-Shi Pang · 2007
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
Lucian Bu, Robert Babu, and Bart De Schutter · 2008
Earlier work this paper cites.
Regret minimization in games with incomplete information
Martin Zinkevich, Michael Johanson, Michael Bowling, and Carmelo Piccione · 2008
Earlier work this paper cites.
Near-optimal no-regret algorithms for zero-sum games
Constantinos Daskalakis, Alan Deckelbaum, and Anthony Kim · 2011
Earlier work this paper cites.
Independent reinforcement learners in cooperative Markov games: a survey regarding coordination problems
Laetitia Matignon, Guillaume J Laurent, and Nadine Le Fort-Piat · 2012
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters · 2013
Earlier work this paper cites.
Sample complexity of episodic fixed-horizon reinforcement learning
Christoph Dann and Emma Brunskill · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Cited alongside, same era.
Decentralized q-learning for stochastic teams and games
Gürdal Arslan and Serdar Yüksel · 2017
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Constantinos Daskalakis, Andrew Ilyas, Vasilis Syrgkanis, and Haoyang Zeng · 2017
Cited alongside, same era.
Stabilising experience replay for deep multi-agent reinforcement learning
Jakob Foerster, Nantas Nardelli, Gregory Farquhar, Triantafyllos Afouras, Philip HS Torr, Pushmeet Kohli, and Shimon Whiteson · 2017
Cited alongside, same era.
A survey of learning in multiagent environments: Dealing with non-stationarity
On the sample complexity of the linear quadratic regulator
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu · 2019
Later among the works it cites.
An accelerated inexact proximal point method for solving nonconvex-concave min-max problems
Weiwei Kong and Renato DC Monteiro · 2019
Later among the works it cites.
Interaction matters: A note on non-asymptotic local convergence of generative adversarial networks
Tengyuan Liang and James Stokes · 2019
Later among the works it cites.
Computing approximate equilibria in sequential adversarial games by exploitability descent
Edward Lockhart, Marc Lanctot, Julien Pérolat, Jean-Baptiste Lespiau, Dustin Morrill, Finbarr Timbers, and Karl Tuyls · 2019
Later among the works it cites.
Optimistic mirror descent in saddle-point problems: Going the extra (gradient) mile
Panayotis Mertikopoulos, Houssam Zenati, Bruno Lecouat, Chuan-Sheng Foo, Vijay Chandrasekhar, and Georgios Piliouras · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pablo Hernandez-Leal, Michael Kaisers, Tim Baarslag, and Enrique Munoz de Cote · 2017
Cited alongside, same era.
Gans trained by a two time-scale update rule converge to a local Nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
Online reinforcement learning in stochastic games
Chen-Yu Wei, Yi-Te Hong, and Chi-Jen Lu · 2017
Cited alongside, same era.
The limit points of (optimistic) gradient descent in min-max optimization
Constantinos Daskalakis and Ioannis Panageas · 2018
Cited alongside, same era.
Is Q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Cited alongside, same era.
Later among the works it cites.
Solving a class of non-convex min-max games using iterative first order methods
Maher Nouiehed, Maziar Sanjabi, Tianjian Huang, Jason D Lee, and Meisam Razaviyayn · 2019
Later among the works it cites.
Shayegan Omidshafiei, Daniel Hennes, Dustin Morrill, Remi Munos, Julien Perolat, Marc Lanctot, Audrunas Gruslys, Jean-Baptiste Lespiau, and Karl Tuyls · 2019
Later among the works it cites.
Efficient algorithms for smooth minimax optimization
Kiran K Thekumparampil, Prateek Jain, Praneeth Netrapalli, and Sewoong Oh · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Later among the works it cites.
Optimality and approximation with policy gradient methods in markov decision processes
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2020
Later among the works it cites.
A tight and unified analysis of extragradient for a whole spectrum of differentiable games
Waïss Azizian, Ioannis Mitliagkas, Simon Lacoste-Julien, and Gauthier Gidel · 2020
Later among the works it cites.
Provable self-play algorithms for competitive reinforcement learning
Yu Bai and Chi Jin · 2020
Later among the works it cites.
Near-optimal reinforcement learning with self-play
Yu Bai, Chi Jin, and Tiancheng Yu · 2020
Later among the works it cites.
Optimistic policy optimization with bandit feedback
Yonathan Efroni, Lior Shani, Aviv Rosenberg, and Shie Mannor · 2020
Later among the works it cites.
A theoretical analysis of deep q-learning
Jianqing Fan, Zhaoran Wang, Yuchen Xie, and Zhuoran Yang · 2020
Later among the works it cites.
Last iterate is slower than averaged iterate in smooth convex-concave saddle point problems
Noah Golowich, Sarath Pattathil, Constantinos Daskalakis, and Asuman Ozdaglar · 2020
Later among the works it cites.
Linear last-iterate convergence for matrix games and stochastic games
Chung-Wei Lee, Haipeng Luo, Chen-Yu Wei, and Mengxiao Zhang · 2020
Later among the works it cites.
On gradient descent ascent for nonconvex-concave minimax problems
Tianyi Lin, Chi Jin, and Michael Jordan · 2020
Later among the works it cites.
Hybrid block successive approximation for one-sided non-convex min-max problems: Algorithms and applications
S. Lu, I. Tsaknakis, M. Hong, and Y. Chen · 2020
Later among the works it cites.
Policy-gradient algorithms have no guarantees of convergence in linear quadratic games
Eric Mazumdar, Lillian J. Ratliff, Michael I. Jordan, and S. Shankar Sastry · 2020
Later among the works it cites.
A unified analysis of extra-gradient and optimistic gradient methods for saddle point problems: Proximal point approach
Aryan Mokhtari, Asuman Ozdaglar, and Sarath Pattathil · 2020
Later among the works it cites.
Learning zero-sum simultaneous-move Markov games using function approximation and correlated equilibrium
Qiaomin Xie, Yudong Chen, Zhaoran Wang, and Zhuoran Yang · 2020
Later among the works it cites.
Global convergence and variance reduction for a class of nonconvex-nonconcave minimax problems
Junchi Yang, Negar Kiyavash, and Niao He · 2020
Later among the works it cites.
Model-based multi-agent RL in zero-sum markov games with near-optimal sample complexity
Kaiqing Zhang, Sham M. Kakade, Tamer Basar, and Lin F. Yang · 2020
Later among the works it cites.