Fetching the paper…
Reading the bibliography…
Multi-agent reinforcement learning (MARL) algorithms often suffer from an exponential sample complexity dependence on the number of agents, a phenomenon known as \emph{the curse of multiagents}.
Stochastic games
Lloyd S Shapley · 1953
Earlier work this paper cites.
Team decision theory and information structures
Yu-Chi Ho · 1980
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
Planning, learning and coordination in multiagent decision processes
Craig Boutilier · 1996
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Caroline Claus and Craig Boutilier · 1998
Earlier work this paper cites.
The Theory of Learning in Games , volume 2
Drew Fudenberg, Fudenberg Drew, David K Levine, and David K Levine · 1998
Earlier work this paper cites.
Adaptive game playing using multiplicative weights
Yoav Freund and Robert E Schapire · 1999
Earlier work this paper cites.
A simple adaptive procedure leading to correlated equilibrium
Sergiu Hart and Andreu Mas-Colell · 2000
Earlier work this paper cites.
An algorithm for distributed reinforcement learning in cooperative multi-agent systems
Martin Lauer and Martin Riedmiller · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Friend-or-Foe Q-learning in general-sum games
Michael L Littman · 2001
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Learning to reach the Pareto optimal Nash equilibrium as a team
Katja Verbeeck, Ann Nowé, Tom Lenaerts, and Johan Parent · 2002
Earlier work this paper cites.
Reinforcement learning to play an optimal Nash equilibrium in team Markov games
Xiaofeng Wang and Tuomas Sandholm · 2002
Earlier work this paper cites.
Uncoupled dynamics do not lead to Nash equilibrium
Sergiu Hart and Andreu Mas-Colell · 2003
Earlier work this paper cites.
Nash Q-learning for general-sum stochastic games
Junling Hu and Michael P Wellman · 2003
Earlier work this paper cites.
Individual Q-learning in normal form games
David S Leslie and Edmund J Collins · 2005
Earlier work this paper cites.
Prediction, Learning, and Games
Nicolo Cesa-Bianchi and Gábor Lugosi · 2006
Earlier work this paper cites.
From external to internal regret
Avrim Blum and Yishay Mansour · 2007
Earlier work this paper cites.
Algorithmic Game Theory
Noam Nisan, Tim Roughgarden, Eva Tardos, and Vijay V Vazirani · 2007
Earlier work this paper cites.
Improved memory-bounded dynamic programming for decentralized POMDPs
Sven Seuken and Shlomo Zilberstein · 2007
Earlier work this paper cites.
Medium access in cognitive radio networks: A competitive multi-armed bandit framework
Lifeng Lai, Hai Jiang, and H Vincent Poor · 2008
Earlier work this paper cites.
Optimal and approximate Q-value functions for decentralized POMDPs
Frans A Oliehoek, Matthijs TJ Spaan, and Nikos Vlassis · 2008
Earlier work this paper cites.
Policy iteration for decentralized control of Markov decision processes
Daniel S Bernstein, Christopher Amato, Eric A Hansen, and Shlomo Zilberstein · 2009
Earlier work this paper cites.
The complexity of computing a Nash equilibrium
Constantinos Daskalakis, Paul W Goldberg, and Christos H Papadimitriou · 2009
Earlier work this paper cites.
Multiplicative updates outperform generic no-regret learning in congestion games
Robert Kleinberg, Georgios Piliouras, and Éva Tardos · 2009
Earlier work this paper cites.
Intrinsic robustness of the price of anarchy
Tim Roughgarden · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Cited alongside, same era.
Strategy iteration is strongly polynomial for 2-player turn-based stochastic games with a constant discount factor
Thomas Dueholm Hansen, Peter Bro Miltersen, and Uri Zwick · 2013
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Cited alongside, same era.
QD-learning: A collaborative distributed strategy for multi-agent reinforcement learning through consensus + innovations
Soummya Kar, José M. F. Moura, and H. Vincent Poor · 2013
Cited alongside, same era.
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters · 2013
Cited alongside, same era.
Composable and efficient mechanisms
Lower bounds for non-convex stochastic optimization
Yossi Arjevani, Yair Carmon, John C Duchi, Dylan J Foster, Nathan Srebro, and Blake Woodworth · 2019
Later among the works it cites.
Momentum-based variance reduction in non-convex SGD
Ashok Cutkosky and Francesco Orabona · 2019
Later among the works it cites.
Learning to collaborate in Markov decision processes
Goran Radanovic, Rati Devidze, David Parkes, and Adish Singla · 2019
Later among the works it cites.
Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning
Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Earl Hostallero, and Yung Yi · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vasilis Syrgkanis and Eva Tardos · 2013
Cited alongside, same era.
No-regret dynamics and fictitious play
Yannick Viossat and Andriy Zapechelnyuk · 2013
Cited alongside, same era.
Concurrent bandits and cognitive radio networks
Orly Avner and Shie Mannor · 2014
Cited alongside, same era.
Reinforcement learning in decentralized stochastic control systems with partial history sharing
Jalal Arabneydi and Aditya Mahajan · 2015
Cited alongside, same era.
Convex optimization: Algorithms and complexity
Sébastien Bubeck et al · 2015
Cited alongside, same era.
Explore no more: Improved high-probability regret bounds for non-stochastic bandits
Gergely Neu · 2015
Cited alongside, same era.
Fast convergence of regularized learning in games
Vasilis Syrgkanis, Alekh Agarwal, Haipeng Luo, and Robert E Schapire · 2015
Cited alongside, same era.
Bora Yongacoglu, Gürdal Arslan, and Serdar Yüksel · 2019
Later among the works it cites.
Online planning for decentralized stochastic control with partial history sharing
Kaiqing Zhang, Erik Miehling, and Tamer Başar · 2019
Later among the works it cites.
Provable self-play algorithms for competitive reinforcement learning
Yu Bai and Chi Jin · 2020
Later among the works it cites.
Near-optimal reinforcement learning with self-play
Yu Bai, Chi Jin, and Tiancheng Yu · 2020
Later among the works it cites.
Independent policy gradient methods for competitive reinforcement learning
Constantinos Daskalakis, Dylan J Foster, and Noah Golowich · 2020
Later among the works it cites.
A high probability analysis of adaptive SGD with momentum
Xiaoyu Li and Francesco Orabona · 2020
Later among the works it cites.
Solving discounted stochastic two-player games with near-optimal time and sample complexity
Aaron Sidford, Mengdi Wang, Lin Yang, and Yinyu Ye · 2020
Later among the works it cites.
Learning zero-sum simultaneous-move Markov games using function approximation and correlated equilibrium
Qiaomin Xie, Yudong Chen, Zhaoran Wang, and Zhuoran Yang · 2020
Later among the works it cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2021
Closest in time.
Online learning for cooperative multi-player multi-armed bandits
William Chang, Mehdi Jafarnia-Jahromi, and Rahul Jain · 2021
Closest in time.
Provably efficient cooperative multi-agent reinforcement learning with function approximation
Abhimanyu Dubey and Alex Pentland · 2021
Closest in time.
Independent natural policy gradient always converges in Markov potential games
Roy Fox, Stephen McAleer, Will Overman, and Ioannis Panageas · 2021
Closest in time.
Decentralized single-timescale actor-critic on zero-sum two-player stochastic games
Hongyi Guo, Zuyue Fu, Zhuoran Yang, and Zhaoran Wang · 2021
Closest in time.
V-learning–A simple, efficient, decentralized algorithm for multiagent RL
Chi Jin, Qinghua Liu, Yuanhao Wang, and Tiancheng Yu · 2021
Closest in time.
Global convergence of multi-agent policy gradient in Markov potential games
Stefanos Leonardos, Will Overman, Ioannis Panageas, and Georgios Piliouras · 2021
Closest in time.
A sharp analysis of model-based reinforcement learning with self-play
Qinghua Liu, Tiancheng Yu, Yu Bai, and Chi Jin · 2021
Closest in time.
UCB momentum Q-learning: Correcting the bias without forgetting
Pierre Menard, Omar Darwiche Domingues, Xuedong Shang, and Michal Valko · 2021
Closest in time.
Learning in nonzero-sum stochastic games with potentials
David Mguni, Yutong Wu, Yali Du, Yaodong Yang, Ziyi Wang, Minne Li, Ying Wen, Joel Jennings, and Jun Wang · 2021
Closest in time.
Decentralized Q-learning in zero-sum Markov games
Muhammed O Sayin, Kaiqing Zhang, David S Leslie, Tamer Başar, and Asuman Ozdaglar · 2021
Closest in time.
When can we learn general-sum Markov games with a large number of players sample-efficiently?
Ziang Song, Song Mei, and Yu Bai · 2021
Closest in time.
Online learning in unknown Markov games
Yi Tian, Yuanhao Wang, Tiancheng Yu, and Suvrit Sra · 2021
Closest in time.
Last-iterate convergence of decentralized optimistic gradient descent/ascent in infinite-horizon competitive Markov games
Chen-Yu Wei, Chung-Wei Lee, Mengxiao Zhang, and Haipeng Luo · 2021
Closest in time.
Gradient play in multi-agent Markov stochastic games: Stationary points and convergence
Runyu Zhang, Zhaolin Ren, and Na Li · 2021
Closest in time.
Provably efficient policy gradient methods for two-player zero-sum Markov games
Yulai Zhao, Yuandong Tian, Jason D Lee, and Simon S Du · 2021
Closest in time.
Provably efficient reinforcement learning in decentralized general-sum Markov games
Weichao Mao and Tamer Başar · 2022
Closest in time.