Fetching the paper…
Reading the bibliography…
We study Nash equilibria learning of a general-sum stochastic game with an unknown transition probability density function.
Non-cooperative games
J. F. Nash · 1951
Earlier work this paper cites.
Stochastic games
L. S. Shapley · 1953
Earlier work this paper cites.
On a stochastic approximation method
K. L. Chung · 1954
Earlier work this paper cites.
Monotone (nonlinear) operators in Hilbert space
G. J. Minty · 1962
Earlier work this paper cites.
On some nonlinear elliptic differential-functional equations
P. Hartman and G. Stampacchia · 1966
Earlier work this paper cites.
On nonterminating stochastic games
A. Hoffman and R. Karp · 1966
Earlier work this paper cites.
An Introduction to Variational Inequalities and Their Applications
D. Kinderlehrer and G. Stampacchia · 1980
Earlier work this paper cites.
Nonlinear programming and stationary equilibria in stochastic games
J. A. Filar, T. A. Schultz, F. Thuijsman, and O. J. Vrieze · 1991
Earlier work this paper cites.
Q-learning
C. J. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
M. L. Littman · 1994
Earlier work this paper cites.
Optimum bounds for the distributions of martingales in Banach spaces
I. Pinelis · 1994
Earlier work this paper cites.
Multiagent reinforcement learning: theoretical framework and an algorithm
J. Hu and M. P. Wellman · 1998
Earlier work this paper cites.
Actor-Critic Algorithms
V. Konda and J. Tsitsiklis · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, S. P. Singh D. A. McAllester, and Y. Mansour · 1999
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
J. Baxter and P. L. Bartlett · 2001
Cited alongside, same era.
Friend-or-Foe Q-learning in general-sum games
M. L. Littman · 2001
Cited alongside, same era.
Finite-Dimensional Variational Inequalities and Complementarity Problems, Vol. I
F. Facchinei and J. S. Pang · 2003
Cited alongside, same era.
Correlated Q-learning
A. Greenwald, K. Hall, and R. Serrano · 2003
Cited alongside, same era.
Nash Q-learning for general-sum stochastic games
J. Hu and M. P. Wellman · 2003
Cited alongside, same era.
A note on approximate Nash equilibria
C. Daskalakis, A. Mehta, and C. Papadimitriou · 2006
Cited alongside, same era.
Optimality and approximation with policy gradient methods in Markov decision processes
A. Agarwal, S. M. Kakade, J. D. Lee, and G. Mahajan · 2020
Later among the works it cites.
Independent policy gradient methods for competitive reinforcement learning
C. Daskalakis, D. J. Foster, and N. Golowich · 2020
Later among the works it cites.
An improved analysis of (variance-reduced) policy gradient and natural policy gradient methods
Y. Liu, K. Zhang, T. Başar, and W. Yin · 2020
Later among the works it cites.
Optimistic dual extrapolation for coherent non-monotone variational inequalities
C. Song, Z. Zhou, Y. Jiang Y. Zhou, and Y. Ma · 2020
Later among the works it cites.
Global convergence of policy gradient methods to (almost) locally optimal policies
K. Zhang, A. Koppel, H. Zhu, and T. Başar · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Koshal, A. Nedić, and U. V. Shanbhag · 2010
Cited alongside, same era.
Competitive Markov Decision Processes
J. A. Filar and K. Vrieze · 2012
Cited alongside, same era.
Regularized iterative stochastic approximation methods for stochastic variational inequality problems
J. Koshal, A. Nedić, and U. V. Shanbhag · 2013
Cited alongside, same era.
Two-timescale algorithms for learning Nash equilibria in general-sum stochastic games
H. L. Prasad, L. A. Prashanth, and S. Bhatnagar · 2015
Cited alongside, same era.
Actor-Critic fictitious play in simultaneous move multistage games
J. Perolat, B. Piot, and O. Pietquin · 2018
Cited alongside, same era.
A finite sample analysis of the Actor-Critic algorithm
Z. Yang, M. Hong K. Zhang, and T. Başar · 2018
Cited alongside, same era.
T. Y. Chen, K. Zhang, G. B. Giannakis, and T. Başar · 2021
Later among the works it cites.
Global convergence of multi-agent policy gradient in Markov potential games
S. Leonardos, W. Overman, I. Panageas, and G. Piliouras · 2021
Later among the works it cites.
First-order convergence theory for weakly-convex-weakly-concave min-max problems
M. Liu, H. Rafique, Q. Lin, and T. Yang · 2021
Later among the works it cites.
Decentralized policy gradient descent ascent for safe multi-agent reinforcement learning
S. Lu, K. Zhang, T. Chen, T. Başar, and L. Horesh · 2021
Later among the works it cites.
Last-iterate convergence of decentralized optimistic gradient descent/ascent in infinite-horizon competitive Markov games
C. Y. Wei, C. W. Lee, M. X. Zhang, and H. P. Luo · 2021
Later among the works it cites.
Gradient play in stochastic games: stationary points, convergence, and sample complexity
R. Zhang, Z. Ren, and N. Li · 2021
Later among the works it cites.
Provably efficient reinforcement learning in decentralized generalsum Markov games
W. C. Mao and T. Başar · 2022
Closest in time.
Provably efficient policy gradient methods for two-player zero-sum Markov games
Y. Zhao, Y. Tian, J. D. Lee, and S. S. Du · 2022
Closest in time.