Fetching the paper…
Reading the bibliography…
Despite the success of single-agent reinforcement learning, multi-agent reinforcement learning (MARL) remains challenging due to complex interactions between agents.
Distributed asynchronous deterministic and stochastic gradient optimization algorithms
J. Tsitsiklis, D. Bertsekas, and M. Athans · 1986
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
M. L. Littman · 1994
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
An algorithm for distributed reinforcement learning in cooperative multi-agent systems
M. Lauer and M. Riedmiller · 2000
Earlier work this paper cites.
Value-function reinforcement learning in Markov games
M. L. Littman · 2001
Earlier work this paper cites.
Nash Q-learning for general-sum stochastic games
J. Hu and M. P. Wellman · 2003
Earlier work this paper cites.
Least squares policy evaluation algorithms with linear function approximation
A. Nedić and D. P. Bertsekas · 2003
Earlier work this paper cites.
Reinforcement learning to play an optimal Nash equilibrium in team Markov games
X. Wang and T. Sandholm · 2003
Earlier work this paper cites.
Coverage control for mobile sensing networks
J. Cortes, S. Martinez, T. Karatas, and F. Bullo · 2004
Earlier work this paper cites.
Distributed optimization in sensor networks
M. Rabbat and R. Nowak · 2004
Earlier work this paper cites.
Networked robots: Flying robot navigation using a sensor net
P. Corke, R. Peterson, and D. Rus · 2005
Earlier work this paper cites.
Multi-task reinforcement learning: A hierarchical Bayesian approach
A. Wilson, A. Fern, S. Ray, and P. Tadepalli · 2007
Earlier work this paper cites.
Stochastic approximation: A dynamical systems viewpoint
V. S. Borkar · 2008
Earlier work this paper cites.
Distributed subgradient methods for multi-agent optimization
A. Nedic and A. Ozdaglar · 2009
Earlier work this paper cites.
A convergent o ( n ) o(n) temporal-difference algorithm for off-policy learning with linear function approximation
R. S. Sutton, H. R. Maei, and C. Szepesvári · 2009
Earlier work this paper cites.
Discrete-time dynamic average consensus
M. Zhu and S. Martínez · 2010
Earlier work this paper cites.
Achieving controllability of electric loads
D. S. Callaway and I. A. Hiskens · 2011
Earlier work this paper cites.
Differentially private empirical risk minimization
K. Chaudhuri, C. Monteleoni, and A. D. Sarwate · 2011
Earlier work this paper cites.
Diffusion adaptation strategies for distributed optimization and learning over networks
J. Chen and A. H. Sayed · 2012
Cited alongside, same era.
Reinforcement learning in robotics: A survey
J. Kober and J. Peters · 2012
Cited alongside, same era.
Distributed optimal power flow for smart microgrids
E. Dall’Anese, H. Zhu, and G. B. Giannakis · 2013
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Cited alongside, same era.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
A. Defazio, F. Bach, and S. Lacoste-Julien · 2014
Cited alongside, same era.
A delayed proximal gradient method with linear convergence rate
H. R. Feyzmahdavian, A. Aytekin, and M. Johansson · 2014
Stochastic variance reduction methods for saddle-point problems
B. Palaniappan and F. Bach · 2016
Later among the works it cites.
Decentralized Q-learning for stochastic teams and games
G. Arslan and S. Yüksel · 2017
Later among the works it cites.
Stochastic variance reduction methods for policy evaluation
S. S. Du, J. Chen, L. Li, L. Xiao, and D. Zhou · 2017
Later among the works it cites.
Stabilising experience replay for deep multi-agent reinforcement learning
J. Foerster, N. Nardelli, G. Farquhar, P. Torr, P. Kohli, S. Whiteson, et al · 2017
Later among the works it cites.
Cooperative multi-agent control using deep reinforcement learning
J. K. Gupta, M. Egorov, and M. Kochenderfer · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Decentralized and privacy-preserving low-rank matrix completion
A. Lin and Q. Ling · 2014
Cited alongside, same era.
Finite-sample analysis of proximal gradient td algorithms
B. Liu, J. Liu, M. Ghavamzadeh, S. Mahadevan, and M. Petrik · 2015
Cited alongside, same era.
Distributed policy evaluation under multiple behavior strategies
S. V. Macua, J. Chen, S. Zazo, and A. H. Sayed · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Actor-mimic: Deep multi-task and transfer reinforcement learning
E. Parisotto, J. L. Ba, and R. Salakhutdinov · 2015
Cited alongside, same era.
Extra: An exact first-order algorithm for decentralized consensus optimization
W. Shi, Q. Ling, G. Wu, and W. Yin · 2015
Cited alongside, same era.
On the convergence rate of incremental aggregated gradient algorithms
M. Gurbuzbalaban, A. Ozdaglar, and P. A. Parrilo · 2017
Later among the works it cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
R. Lowe, Y. Wu, A. Tamar, J. Harb, P. Abbeel, and I. Mordatch · 2017
Later among the works it cites.
Diff-dac: Distributed actor-critic for multitask deep reinforcement learning
S. V. Macua, A. Tukiainen, D. G.-O. Hernández, D. Baldazo, E. M. de Cote, and S. Zazo · 2017
Later among the works it cites.
Deep decentralized multi-task multi-agent reinforcement learning under partial observability
S. Omidshafiei, J. Pazis, C. Amato, J. P. How, and J. Vian · 2017
Later among the works it cites.
Harnessing smoothness to accelerate distributed optimization
G. Qu and N. Li · 2017
Later among the works it cites.
Minimizing finite sums with the stochastic average gradient
M. Schmidt, N. Le Roux, and F. Bach · 2017
Later among the works it cites.
Mastering the game of Go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Later among the works it cites.
Distral: Robust multi-task reinforcement learning
Y. W. Teh, V. Bapst, W. M. Czarnecki, J. Quan, J. Kirkpatrick, R. Hadsell, N. Heess, and R. Pascanu · 2017
Later among the works it cites.
M. Wang · 2017
Later among the works it cites.
Primal-dual algorithm for distributed reinforcement learning: Distributed gtd2
D. Lee, H. Yoon, and N. Hovakimyan · 2018
Closest in time.
Distributed stochastic gradient tracking methods
S. Pu and A. Nedić · 2018
Closest in time.
Fully decentralized multi-agent reinforcement learning with networked agents
K. Zhang, Z. Yang, H. Liu, T. Zhang, and T. Başar · 2018
Closest in time.