Fetching the paper…
Reading the bibliography…
We examine global non-asymptotic convergence properties of policy gradient methods for multi-agent reinforcement learning (RL) problems in Markov potential games (MPG).
An iterative method of solving a game
Robinson, J · 1951
Earlier work this paper cites.
Stochastic games
Shapley, L. S · 1953
Earlier work this paper cites.
Stochastic games with infinitely many strategies
Takahashi, M · 1962
Earlier work this paper cites.
Equilibrium in a stochastic n n -person game
Fink, A. M · 1964
Earlier work this paper cites.
On stochastic games
Maitra, A. and Parthasarathy, T · 1970
Earlier work this paper cites.
On stochastic games, II
Maitra, A. and Parthasarathy, T · 1971
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Fictitious play property for games with identical interests
Monderer, D. and Shapley, L · 1996
Earlier work this paper cites.
Contraction conditions for average and α \alpha -discount optimality in countable state Markov games with unbounded rewards
Altman, E., Hordijk, A., and Spieksma, F · 1997
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M · 2001
Earlier work this paper cites.
On the global convergence of stochastic fictitious play
Hofbauer, J. and Sandholm, W. H · 2002
Earlier work this paper cites.
Reinforcement learning to play an optimal Nash equilibrium in team Markov games
Wang, X. and Sandholm, T · 2002
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Zinkevich, M · 2003
Earlier work this paper cites.
Prediction, Learning, and Games
Cesa-Bianchi, N. and Lugosi, G · 2006
Earlier work this paper cites.
The stochastic lake game: A numerical solution
Dechert, W. D. and O’Donnell, S · 2006
Earlier work this paper cites.
Lecture 6 in toward theoretical understanding of deep learning, October 2008
Arora, S · 2008
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
Busoniu, L., Babuska, R., and De Schutter, B · 2008
Earlier work this paper cites.
Risk bounds for linear regression, 2009
Audibert, J.-Y. and Catoni, O · 2009
Earlier work this paper cites.
Multiplicative updates outperform generic no-regret learning in congestion games
Kleinberg, R., Piliouras, G., and Tardos, É · 2009
Earlier work this paper cites.
Random design analysis of ridge regression
Hsu, D., Kakade, S. M., and Zhang, T · 2012
Earlier work this paper cites.
State based potential games
Marden, J. R · 2012
Earlier work this paper cites.
Independent reinforcement learners in cooperative Markov games: A survey regarding coordination problems
Matignon, L., Laurent, G. J., and Le Fort-Piat, N · 2012
Earlier work this paper cites.
Discrete–time stochastic control and dynamic potential games: the Euler–Equation approach
González-Sánchez, D. and Hernández-Lerma, O · 2013
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S. and Ben-David, S · 2014
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P · 2016
Earlier work this paper cites.
Safe, multi-agent, reinforcement learning for autonomous driving
Shalev-Shwartz, S., Shammah, S., and Shashua, A · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Dynamic potential games with constraints: Fundamentals and applications in communications
Zazo, S., Macua, S. V., Sánchez-Fernández, M., and Zazo, J · 2016
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R., WU, Y., Tamar, A., Harb, J., Pieter Abbeel, O., and Mordatch, I · 2017
Cited alongside, same era.
Linear-quadratic discrete-time dynamic potential games
Mazalov, V. V., Rettieva, A. N., and Avrachenkov, K. E · 2017
Cited alongside, same era.
Multiplicative weights update with constant step-size in congestion games: Convergence, limit cycles and chaos
Palaiopanos, G., Panageas, I., and Piliouras, G · 2017
Cited alongside, same era.
Multiplicative weights update in zero-sum games
Bailey, J. P. and Piliouras, G · 2018
Cited alongside, same era.
The limit points of (optimistic) gradient descent in min-max optimization
Fast policy extragradient methods for competitive games with entropy regularization
Cen, S., Wei, Y., and Chi, Y · 2021
Later among the works it cites.
Cooperative AI: machines must learn to find common ground, 2021
Dafoe, A., Bachrach, Y., Hadfield, G., Horvitz, E., Larson, K., and Graepel, T · 2021
Later among the works it cites.
Provably efficient cooperative multi-agent reinforcement learning with function approximation
Dubey, A. and Pentland, A · 2021
Later among the works it cites.
Policy gradient methods find the Nash equilibrium in N-player general-sum linear-quadratic games
Hambly, B. M., Xu, R., and Yang, H · 2021
Later among the works it cites.
Exploration-exploitation in multi-agent competition: Convergence with bounded rationality
Leonardos, S., Piliouras, G., and Spendlove, K · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daskalakis, C. and Panageas, I · 2018
Cited alongside, same era.
Learning parametric closed-loop policies for Markov potential games
Macua, S. V., Zazo, J., and Zazo, S · 2018
Cited alongside, same era.
Decentralised learning in systems with many, many strategic agents
Mguni, D., Jennings, J., and de Cote, E. M · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al · 2018
Cited alongside, same era.
Politex: Regret bounds for policy iteration using expert prediction
Abbasi-Yadkori, Y., Bartlett, P., Bhatia, K., Lazic, N., Szepesvari, C., and Weisz, G · 2019
Cited alongside, same era.
Fast and furious learning in zero-sum games: Vanishing regret with non-vanishing step sizes
Bailey, J. and Piliouras, G · 2019
Cited alongside, same era.
Global optimality guarantees for policy gradient methods
Bhandari, J. and Russo, D · 2019
Cited alongside, same era.
A sharp analysis of model-based reinforcement learning with self-play
Liu, Q., Yu, T., Bai, Y., and Jin, C · 2021
Later among the works it cites.
Policy optimization in adversarial MDPs: Improved exploration via dilated bonuses
Luo, H., Wei, C.-Y., and Lee, C.-W · 2021
Later among the works it cites.
Learning in nonzero-sum stochastic games with potentials
Mguni, D. H., Wu, Y., Du, Y., Yang, Y., Wang, Z., Li, M., Wen, Y., Jennings, J., and Wang, J · 2021
Later among the works it cites.
Independent learning in stochastic games
Ozdaglar, A., Sayin, M. O., and Zhang, K · 2021
Later among the works it cites.
FACMAC: Factored multi-agent centralised policy gradients
Peng, B., Rashid, T., Schroeder de Witt, C., Kamienny, P.-A., Torr, P., Böhmer, W., and Whiteson, S · 2021
Later among the works it cites.
Decentralized Q-learning in zero-sum Markov games
Sayin, M. O., Zhang, K., Leslie, D. S., Basar, T., and Ozdaglar, A. E · 2021
Later among the works it cites.
Normative disagreement as a challenge for cooperative AI
Stastny, J., Riché, M., Lyzhov, A., Treutlein, J., Dafoe, A., and Clifton, J · 2021
Later among the works it cites.
Trott, A., Srinivasa, S., van der Wal, D., Haneuse, S., and Zheng, S · 2021
Later among the works it cites.
Exponential lower bounds for planning in MDPs with linearly-realizable optimal action-value functions
Weisz, G., Amortila, P., and Szepesvári, C · 2021
Later among the works it cites.
The surprising effectiveness of PPO in cooperative, multi-agent games
Yu, C., Velu, A., Vinitsky, E., Wang, Y., Bayen, A., and Wu, Y · 2021
Later among the works it cites.
Zhan, W., Cen, S., Huang, B., Chen, Y., Lee, J. D., and Chi, Y · 2021
Later among the works it cites.
Provably efficient policy gradient methods for two-player zero-sum Markov games
Zhao, Y., Tian, Y., Lee, J. D., and Du, S. S · 2021
Later among the works it cites.
Independent natural policy gradient always converges in Markov potential games
Fox, R., Mcaleer, S. M., Overman, W., and Panageas, I · 2022
Closest in time.
Towards general function approximation in zero-sum Markov games
Huang, B., Lee, J. D., Wang, Z., and Yang, Z · 2022
Closest in time.
Decentralized cooperative reinforcement learning with hierarchical information structure
Kao, H., Wei, C.-Y., and Subramanian, V · 2022
Closest in time.
Policy mirror descent for reinforcement learning: Linear convergence, new sampling complexity, and generalized problem classes
Lan, G · 2022
Closest in time.
Exploration-exploitation in multi-agent learning: Catastrophe theory meets game theory
Leonardos, S. and Piliouras, G · 2022
Closest in time.
Global convergence of multi-agent policy gradient in Markov potential games
Leonardos, S., Overman, W., Panageas, I., and Piliouras, G · 2022
Closest in time.
On improving model-free algorithms for decentralized multi-agent reinforcement learning
Mao, W., Basar, T., Yang, L. F., and Zhang, K · 2022
Closest in time.
When can we learn general-sum Markov games with a large number of players sample-efficiently?
Song, Z., Mei, S., and Bai, Y · 2022
Closest in time.
On the convergence rates of policy gradient methods
Xiao, L · 2022
Closest in time.
Zhang, R., Mei, J., Dai, B., Schuurmans, D., and Li, N · 2022
Closest in time.