Fetching the paper…
Reading the bibliography…
A major challenge in multi-agent systems is that the system complexity grows dramatically with the number of agents as well as the size of their action spaces, which is typical in real world scenarios such as autonomous vehicles, robotic teams, network routing, etc.
Non-cooperative games
J. Nash · 1951
Earlier work this paper cites.
Evolution, learning and economic behavior
R. Selten · 1989
Earlier work this paper cites.
The weighted majority algorithm
N. Littlestone and M. K. Warmuth · 1994
Earlier work this paper cites.
Quantal response equilibria for normal form games
R. D. McKelvey and T. R. Palfrey · 1995
Earlier work this paper cites.
Fictitious play property for games with identical interests
D. Monderer and L. S. Shapley · 1996
Earlier work this paper cites.
Potential games
D. Monderer and L. S. Shapley · 1996
Earlier work this paper cites.
The theory of probability
H. Jeffreys · 1998
Earlier work this paper cites.
Adaptive game playing using multiplicative weights
Y. Freund and R. E. Schapire · 1999
Earlier work this paper cites.
A natural policy gradient
S. M. Kakade · 2001
Earlier work this paper cites.
Strategic learning and its limits
H. P. Young · 2004
Earlier work this paper cites.
Convergence and approximation in potential games
G. Christodoulou, V. S. Mirrokni, and A. Sidiropoulos · 2006
Earlier work this paper cites.
Math 280 (probability theory) lecture notes, April 2007
B. K. Driver · 2007
Earlier work this paper cites.
Regret based dynamics: convergence in weakly acyclic games
J. R. Marden, G. Arslan, and J. S. Shamma · 2007
Earlier work this paper cites.
Inapproximability of pure Nash equilibria
A. Skopalik and B. Vöcking · 2008
Earlier work this paper cites.
Payoff-based dynamics for multiplayer weakly acyclic games
J. R. Marden, H. P. Young, G. Arslan, and J. S. Shamma · 2009
Earlier work this paper cites.
Convergence to approximate nash equilibria in congestion games
S. Chien and A. Sinclair · 2011
Earlier work this paper cites.
The multiplicative weights update method: a meta-algorithm and applications
S. Arora, E. Hazan, and S. Kale · 2012
Earlier work this paper cites.
On the complexity of approximating a Nash equilibrium
C. Daskalakis · 2013
Earlier work this paper cites.
Decentralized q-learning for stochastic teams and games
G. Arslan and S. Yüksel · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Cited alongside, same era.
Learning in games via reinforcement and regularization
P. Mertikopoulos and W. H. Sandholm · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Cited alongside, same era.
Learning with bandit feedback in potential games
A. Heliou, J. Cohen, and P. Mertikopoulos · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
R. Lowe, Y. I. Wu, A. Tamar, J. Harb, O. Pieter Abbeel, and I. Mordatch · 2017
Cited alongside, same era.
Multiplicative weights update with constant step-size in congestion games: Convergence, limit cycles and chaos
Fast policy extragradient methods for competitive games with entropy regularization
S. Cen, Y. Wei, and Y. Chi · 2021
Later among the works it cites.
Near-optimal no-regret learning in general games
C. Daskalakis, M. Fishelson, and N. Golowich · 2021
Later among the works it cites.
Independent natural policy gradient always converges in Markov potential games
R. Fox, S. McAleer, W. Overman, and I. Panageas · 2021
Later among the works it cites.
V-learning–a simple, efficient, decentralized algorithm for multiagent RL
C. Jin, Q. Liu, Y. Wang, and T. Yu · 2021
Later among the works it cites.
G. Lan · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Palaiopanos, I. Panageas, and G. Piliouras · 2017
Cited alongside, same era.
Analysis of Best Response Dynamics in Potential Games
S. Durand · 2018
Cited alongside, same era.
Global convergence of policy gradient methods for the linear quadratic regulator
M. Fazel, R. Ge, S. Kakade, and M. Mesbahi · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Global optimality guarantees for policy gradient methods
J. Bhandari and D. Russo · 2019
Cited alongside, same era.
Neural policy gradient methods: Global optimality and rates of convergence
L. Wang, Q. Cai, Z. Yang, and Z. Wang · 2019
Cited alongside, same era.
Optimality and approximation with policy gradient methods in Markov decision processes
A. Agarwal, S. M. Kakade, J. D. Lee, and G. Mahajan · 2020
Cited alongside, same era.
Global convergence of multi-agent policy gradient in Markov potential games
S. Leonardos, W. Overman, I. Panageas, and G. Piliouras · 2021
Later among the works it cites.
Softmax policy gradient methods can take exponential time to converge
G. Li, Y. Wei, Y. Chi, Y. Gu, and Y. Chen · 2021
Later among the works it cites.
A graph placement methodology for fast chip design
A. Mirhoseini, A. Goldie, M. Yazgan, J. W. Jiang, E. Songhori, S. Wang, Y.-J. Lee, E. Johnson, O. Pathak, A. Nazi, et al · 2021
Later among the works it cites.
When can we learn general-sum Markov games with a large number of players sample-efficiently?
Z. Song, S. Mei, and Y. Bai · 2021
Later among the works it cites.
C.-Y. Wei, C.-W. Lee, M. Zhang, and H. Luo · 2021
Later among the works it cites.
W. Zhan, S. Cen, B. Huang, Y. Chen, J. D. Lee, and Y. Chi · 2021
Later among the works it cites.
Gradient play in stochastic games: stationary points, convergence, and sample complexity
R. Zhang, Z. Ren, and N. Li · 2021
Later among the works it cites.
Provably efficient policy gradient methods for two-player zero-sum Markov games
Y. Zhao, Y. Tian, J. D. Lee, and S. S. Du · 2021
Later among the works it cites.
D. Ding, C.-Y. Wei, K. Zhang, and M. R. Jovanović · 2022
Closest in time.
Provably efficient reinforcement learning in decentralized general-sum markov games
W. Mao and T. Başar · 2022
Closest in time.
On improving model-free algorithms for decentralized multi-agent reinforcement learning
W. Mao, L. Yang, K. Zhang, and T. Basar · 2022
Closest in time.
On the convergence rates of policy gradient methods
L. Xiao · 2022
Closest in time.
R. Zhang, J. Mei, B. Dai, D. Schuurmans, and N. Li · 2022
Closest in time.