Fetching the paper…
Reading the bibliography…
Multi-Agent Reinforcement Learning (MARL) -- where multiple agents learn to interact in a shared dynamic environment -- permeates across a wide range of critical applications.
Stochastic games
L. S. Shapley · 1953
Earlier work this paper cites.
Discounted markov games: Generalized policy iteration method
J. Van Der Wal · 1978
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
R. J. Williams and J. Peng · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Quantal response equilibria for normal form games
R. D. McKelvey and T. R. Palfrey · 1995
Earlier work this paper cites.
Adaptive game playing using multiplicative weights
Y. Freund and R. E. Schapire · 1999
Earlier work this paper cites.
Stochastic shortest path games
S. D. Patek and D. P. Bertsekas · 1999
Earlier work this paper cites.
Actor-critic algorithms
V. R. Konda and J. N. Tsitsiklis · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
A natural policy gradient
S. M. Kakade · 2002
Earlier work this paper cites.
Natural actor-critic
J. Peters and S. Schaal · 2008
Earlier work this paper cites.
Near-optimal no-regret algorithms for zero-sum games
C. Daskalakis, A. Deckelbaum, and A. Kim · 2011
Earlier work this paper cites.
Optimization, learning, and games with predictable sequences
A. Rakhlin and K. Sridharan · 2013
Earlier work this paper cites.
Approximate dynamic programming for two-player zero-sum Markov games
J. Perolat, B. Scherrer, B. Piot, and O. Pietquin · 2015
Earlier work this paper cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Multiplicative weights update in zero-sum games
J. P. Bailey and G. Piliouras · 2018
Cited alongside, same era.
Last-iterate convergence: Zero-sum games and constrained min-max optimization
C. Daskalakis and I. Panageas · 2018
Cited alongside, same era.
Global optimality guarantees for policy gradient methods
J. Bhandari and D. Russo · 2019
Cited alongside, same era.
A theory of regularized Markov decision processes
M. Geist, B. Scherrer, and O. Pietquin · 2019
Cited alongside, same era.
Softmax policy gradient methods can take exponential time to converge
G. Li, Y. Wei, Y. Chi, Y. Gu, and Y. Chen · 2021
Later among the works it cites.
A sharp analysis of model-based reinforcement learning with self-play
Q. Liu, T. Yu, Y. Bai, and C. Jin · 2021
Later among the works it cites.
Decentralized Q-learning in zero-sum Markov games
M. Sayin, K. Zhang, D. Leslie, T. Basar, and A. Ozdaglar · 2021
Later among the works it cites.
Last-iterate convergence of decentralized optimistic gradient descent/ascent in infinite-horizon competitive markov games
C.-Y. Wei, C.-W. Lee, M. Zhang, and H. Luo · 2021
Later among the works it cites.
W. Zhan, S. Cen, B. Huang, Y. Chen, J. D. Lee, and Y. Chi · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimality and approximation with policy gradient methods in Markov decision processes
A. Agarwal, S. M. Kakade, J. D. Lee, and G. Mahajan · 2020
Cited alongside, same era.
Provable self-play algorithms for competitive reinforcement learning
Y. Bai and C. Jin · 2020
Cited alongside, same era.
Near-optimal reinforcement learning with self-play
Y. Bai, C. Jin, and T. Yu · 2020
Cited alongside, same era.
A note on the linear convergence of policy gradient methods
J. Bhandari and D. Russo · 2020
Cited alongside, same era.
Independent policy gradient methods for competitive reinforcement learning
C. Daskalakis, D. J. Foster, and N. Golowich · 2020
Cited alongside, same era.
On the global convergence rates of softmax policy gradient methods
J. Mei, C. Xiao, C. Szepesvari, and D. Schuurmans · 2020
Cited alongside, same era.
Learning zero-sum simultaneous-move markov games using function approximation and correlated equilibrium
Q. Xie, Y. Chen, Z. Wang, and Z. Yang · 2020
Cited alongside, same era.
A natural actor-critic framework for zero-sum Markov games
A. Alacaoglu, L. Viano, N. He, and V. Cevher · 2022
Closest in time.
Independent natural policy gradient methods for potential games: Finite-time global convergence with entropy regularization
S. Cen, F. Chen, and Y. Chi · 2022
Closest in time.
Policy mirror descent for reinforcement learning: Linear convergence, new sampling complexity, and generalized problem classes
G. Lan · 2022
Closest in time.
Minimax-optimal multi-agent RL in zero-sum Markov games with a generative model
G. Li, Y. Chi, Y. Wei, and Y. Chen · 2022
Closest in time.
S. Sokota, R. D’Orazio, J. Z. Kolter, N. Loizou, M. Lanctot, I. Mitliagkas, N. Brown, and C. Kroer · 2022
Closest in time.
On the convergence rates of policy gradient methods
L. Xiao · 2022
Closest in time.
Y. Yang and C. Ma · 2022
Closest in time.
Regularized gradient descent ascent for two-player zero-sum Markov games
S. Zeng, T. T. Doan, and J. Romberg · 2022
Closest in time.
Policy optimization for Markov games: Unified framework and faster convergence
R. Zhang, Q. Liu, H. Wang, C. Xiong, N. Li, and Y. Bai · 2022
Closest in time.
Provably efficient policy optimization for two-player zero-sum markov games
Y. Zhao, Y. Tian, J. Lee, and S. Du · 2022
Closest in time.