Fetching the paper…
Reading the bibliography…
Non-stationarity is a fundamental challenge in multi-agent reinforcement learning (MARL), where agents update their behaviour as they learn.
Equilibrium in a stochastic n n -person game
Arlington M. Fink · 1964
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Ming Tan · 1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L. Littman · 1994
Earlier work this paper cites.
Asynchronous stochastic approximation and Q-learning
J. N. Tsitsiklis · 1994
Earlier work this paper cites.
A generalized reinforcement-learning model: Convergence and applications
Michael L. Littman and Csaba Szepesvári · 1996
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Caroline Claus and Craig Boutilier · 1998
Earlier work this paper cites.
On the nonconvergence of fictitious play in coordination games
D. P. Foster and H. P. Young · 1998
Earlier work this paper cites.
Actor-critic–type learning algorithms for Markov decision processes
Vijaymohan R Konda and Vivek S Borkar · 1999
Earlier work this paper cites.
Convergence results for single-step on-policy reinforcement-learning algorithms
Satinder Singh, Tommi Jaakkola, Michael L Littman, and Csaba Szepesvári · 2000
Earlier work this paper cites.
Friend-or-foe Q-learning in general-sum games
Michael L. Littman · 2001
Earlier work this paper cites.
Reinforcement learning in Markovian evolutionary games
Vivek S Borkar · 2002
Earlier work this paper cites.
Multiagent learning using a variable learning rate
Michael Bowling and Manuela Veloso · 2002
Earlier work this paper cites.
Nash Q-learning for general-sum stochastic games
Junling Hu and Michael P. Wellman · 2003
Earlier work this paper cites.
Convergent multiple-timescales reinforcement learning algorithms in normal form games
David S Leslie and Edmund J Collins · 2003
Cited alongside, same era.
Stochastic imitation in finite games
Jens Josephson and Alexander Matros · 2004
Cited alongside, same era.
Strategic Learning and its Limits
H. Peyton Young · 2004
Cited alongside, same era.
Individual Q-learning in normal form games
D. Leslie and E. Collins · 2005
Cited alongside, same era.
Regret testing: Learning to play Nash equilibrium without knowing you have an opponent
Dean Foster and H. Peyton Young · 2006
Cited alongside, same era.
Global Nash convergence of Foster and Young’s regret testing
Fabrizio Germano and Gabor Lugosi · 2007
Cited alongside, same era.
A survey and critique of multiagent deep reinforcement learning
Pablo Hernandez-Leal, Bilal Kartal, and Matthew E Taylor · 2019
Later among the works it cites.
Policy-gradient algorithms have no guarantees of convergence in linear quadratic games
Eric Mazumdar, Lillian J Ratliff, Michael I Jordan, and S Shankar Sastry · 2019
Later among the works it cites.
Independent policy gradient methods for competitive reinforcement learning
Constantinos Daskalakis, Dylan J Foster, and Noah Golowich · 2020
Later among the works it cites.
Learning in nonzero-sum stochastic games with potentials
David H Mguni, Yutong Wu, Yali Du, Yaodong Yang, Ziyi Wang, Minne Li, Ying Wen, Joel Jennings, and Jun Wang · 2021
Later among the works it cites.
Independent learning in stochastic games
Asuman Ozdaglar, Muhammed O Sayin, and Kaiqing Zhang · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Payoff-based dynamics for multiplayer weakly acyclic games
Jason R. Marden, H. Peyton Young, Gürdal Arslan, and Jeff S. Shamma · 2009
Cited alongside, same era.
Error bounds for constant step-size Q-learning
Carolyn L Beck and Rayadurgam Srikant · 2012
Cited alongside, same era.
Independent reinforcement learners in cooperative Markov games: a survey regarding coordination problems
Laetitia Matignon, Guillaume J. Laurent, and Nadine Le Fort-Piat · 2012
Cited alongside, same era.
Discounted stochastic games with no stationary Nash equilibrium: two examples
Yehuda Levy · 2013
Cited alongside, same era.
Decentralized Q-learning for stochastic teams and games
Gürdal Arslan and Serdar Yüksel · 2017
Cited alongside, same era.
A survey of learning in multiagent environments: Dealing with non-stationarity
Pablo Hernandez-Leal, Michael Kaisers, Tim Baarslag, and Enrique Munoz de Cote · 2017
Cited alongside, same era.
Muhammed Sayin, Kaiqing Zhang, David Leslie, Tamer Başar, and Asuman Ozdaglar · 2021
Later among the works it cites.
Gradient play in multi-agent Markov stochastic games: Stationary points and convergence
Runyu Zhang, Zhaolin Ren, and Na Li · 2021
Later among the works it cites.
Global convergence of multi-agent policy gradient in Markov potential games
Stefanos Leonardos, Will Overman, Ioannis Panageas, and Georgios Piliouras · 2022
Later among the works it cites.
Independent and decentralized learning in markov potential games
Chinmay Maheshwari, Manxi Wu, Druv Pai, and Shankar Sastry · 2022
Later among the works it cites.
Logit-Q learning in Markov games
Muhammed O Sayin and Onur Unlu · 2022
Later among the works it cites.
Decentralized learning for optimality in stochastic dynamic teams and games with local control and global state information
Bora Yongacoglu, Gürdal Arslan, and Serdar Yüksel · 2022
Later among the works it cites.
Dealing with non-stationarity in decentralized cooperative multi-agent deep reinforcement learning via multi-timescale learning
Hadi Nekoei, Akilesh Badrinaaraayanan, Amit Sinha, Mohammad Amini, Janarthanan Rajendran, Aditya Mahajan, and Sarath Chandar · 2023
Closest in time.