Fetching the paper…
Reading the bibliography…
Recent developments in deep reinforcement learning are concerned with creating decision-making agents which can perform well in various complex domains.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Ming Tan · 1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
Temporal difference learning and TD-gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Rational and convergent learning in stochastic games
Michael Bowling and Manuela Veloso · 2001
Earlier work this paper cites.
Correlated Q-learning
Amy Greenwald, Keith Hall, and Roberto Serrano · 2003
Earlier work this paper cites.
Dealing with non-stationary environments using context detection
Bruno C Da Silva, Eduardo W Basso, Ana LC Bazzan, and Paulo M Engel · 2006
Earlier work this paper cites.
On the role of tracking in stationary environments
Richard S Sutton, Anna Koop, and David Silver · 2007
Earlier work this paper cites.
Multi-agent learning with policy prediction
Chongjie Zhang and Victor Lesser · 2010
Earlier work this paper cites.
Comparative evaluation of MAL algorithms in a diverse set of ad hoc team problems
Stefano V Albrecht and Subramanian Ramamoorthy · 2012
Earlier work this paper cites.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Individual planning in open and typed agent systems
Muthukumaran Chandrasekaran, Adam Eck, Prashant Doshi, and Leenkiat Soh · 2016
Earlier work this paper cites.
Learning to communicate with deep multi-agent reinforcement learning
Jakob Foerster, Ioannis Alexandros Assael, Nando de Freitas, and Shimon Whiteson · 2016
Cited alongside, same era.
Learning to communicate to solve riddles with deep distributed recurrent q-networks
Jakob Foerster, Yannis M Assael, Nando de Freitas, and Shimon Whiteson · 2016
Cited alongside, same era.
Opponent modeling in deep reinforcement learning
He He, Jordan Boyd-Graber, Kevin Kwok, and Hal Daumé III · 2016
Cited alongside, same era.
Learning multiagent communication with backpropagation
Sainbayar Sukhbaatar, Arthur Szlam, and Rob Fergus · 2016
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Stabilising experience replay for deep multi-agent reinforcement learning
Counterfactual multi-agent policy gradients
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2018
Later among the works it cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke van Hoof, and David Meger · 2018
Later among the works it cites.
Learning policy representations in multiagent systems
Aditya Grover, Maruan Al-Shedivat, Jayesh Gupta, Yuri Burda, and Harrison Edwards · 2018
Later among the works it cites.
Emergence of grounded compositional language in multi-agent populations
Igor Mordatch and Pieter Abbeel · 2018
Later among the works it cites.
Machine theory of mind
Neil Rabinowitz, Frank Perbet, Francis Song, Chiyuan Zhang, SM Ali Eslami, and Matthew Botvinick · 2018
Later among the works it cites.
Modeling others using oneself in multi-agent reinforcement learning
Roberta Raileanu, Emily Denton, Arthur Szlam, and Rob Fergus · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jakob Foerster, Nantas Nardelli, Gregory Farquhar, Triantafyllos Afouras, Philip HS Torr, Pushmeet Kohli, and Shimon Whiteson · 2017
Cited alongside, same era.
A survey of learning in multiagent environments: Dealing with non-stationarity
Pablo Hernandez-Leal, Michael Kaisers, Tim Baarslag, and Enrique Munoz de Cote · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2017
Cited alongside, same era.
Continuous adaptation via meta-learning in nonstationary and competitive environments
Maruan Al-Shedivat, Trapit Bansal, Yuri Burda, Ilya Sutskever, Igor Mordatch, and Pieter Abbeel · 2018
Cited alongside, same era.
Autonomous agents modelling other agents: A comprehensive survey and open problems
Stefano V Albrecht and Peter Stone · 2018
Cited alongside, same era.
Later among the works it cites.
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder Witt, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson · 2018
Later among the works it cites.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al · 2018
Later among the works it cites.
A survey on transfer learning for multiagent reinforcement learning systems
Felipe Leno Da Silva and Anna Helena Reali Costa · 2019
Closest in time.
Stable opponent shaping in differentiable games
Alistair Letcher, Jakob Foerster, David Balduzzi, Tim Rocktäschel, and Shimon Whiteson · 2019
Closest in time.
Robust multi-agent reinforcement learning via minimax deep deterministic policy gradient
Shihui Li, Yi Wu, Xinyue Cui, Honghua Dong, Fei Fang, and Stuart Russell · 2019
Closest in time.
Learning when to communicate at scale in multiagent cooperative and competitive tasks
Amanpreet Singh, Tushar Jain, and Sainbayar Sukhbaatar · 2019
Closest in time.
Relational forward models for multi-agent learning
Andrea Tacchetti, H. Francis Song, Pedro A. M. Mediano, Vinicius Zambaldi, János Kramár, Neil C. Rabinowitz, Thore Graepel, Matthew Botvinick, and Peter W. Battaglia · 2019
Closest in time.