Fetching the paper…
Reading the bibliography…
We consider a multi-agent reinforcement learning problem where each agent seeks to maximize a shared reward while interacting with other agents, and they may or may not be able to communicate.
Learning to predict by the methods of temporal differences
Richard S. Sutton · 1988
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Caroline Claus and Craig Boutilier · 1998
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1998
Earlier work this paper cites.
Readings in agents
Ming Tan · 1998
Earlier work this paper cites.
An algorithm for distributed reinforcement learning in cooperative multi-agent systems
Martin Lauer and Martin Riedmiller · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A. McAllester, Satinder P. Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Extending q-learning to general adaptive multi-agent systems
Gerald Tesauro · 2004
Earlier work this paper cites.
Bounded policy iteration for decentralized pomdps
Daniel S Bernstein, Eric A Hansen, and Shlomo Zilberstein · 2005
Earlier work this paper cites.
Hysteretic q-learning :an algorithm for decentralized reinforcement learning in cooperative multi-agent teams
L. Matignon, G. J. Laurent, and N. L. Fort-Piat · 2007
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
Lucian Bu, Robert Babu, Bart De Schutter, et al · 2008
Earlier work this paper cites.
Optimal and approximate q-value functions for decentralized pomdps
Frans A Oliehoek, Matthijs TJ Spaan, and Nikos Vlassis · 2008
Earlier work this paper cites.
Search-based structured prediction
Hal Daumé, John Langford, and Daniel Marcu · 2009
Earlier work this paper cites.
Multi-agent reinforcement learning : An overview
Lucian Busoniu, Robert Babuska, and Bart De Schutter · 2010
Earlier work this paper cites.
Efficient reductions for imitation learning
Stéphane Ross and Drew Bagnell · 2010
Earlier work this paper cites.
No-regret reductions for imitation learning and structured prediction
Stéphane Ross, Geoffrey J. Gordon, and J. Andrew Bagnell · 2010
Earlier work this paper cites.
Review: Independent reinforcement learners in cooperative Markov games: A survey regarding coordination problems
Laetitia Matignon, Guillaume j. Laurent, and Nadine Le Fort-Piat · 2012
Earlier work this paper cites.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael I. Jordan, and Pieter Abbeel · 2015
Cited alongside, same era.
Multi-agent deep reinforcement learning, 2016
Maxim Egorov · 2016
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Emergence of grounded compositional language in multi-agent populations
Igor Mordatch and Pieter Abbeel · 2017
Later among the works it cites.
Deep decentralized multi-task multi-agent reinforcement learning under partial observability
Shayegan Omidshafiei, Jason Pazis, Christopher Amato, Jonathan P. How, and John Vian · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Later among the works it cites.
Value-decomposition networks for cooperative multi-agent learning
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Multi-agent reinforcement learning as a rehearsal for decentralized planning
Landon Kraemer and Bikramjit Banerjee · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Cited alongside, same era.
The concrete distribution: A continuous relaxation of discrete random variables
Chris J Maddison, Andriy Mnih, and Yee Whye Teh · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Nicolas Usunier, Gabriel Synnaeve, Zeming Lin, and Soumith Chintala · 2016
Cited alongside, same era.
Fully decentralized policies for multi-agent systems: An information theoretic approach
Roel Dobbe, David Fridovich-Keil, and Claire Tomlin · 2017
Cited alongside, same era.
Jakob N. Foerster, Christian A. Schröder de Witt, Gregory Farquhar, Philip H. S. Torr, Wendelin Boehmer, and Shimon Whiteson · 2018
Later among the works it cites.
An algorithmic perspective on imitation learning
Takayuki Osa, Joni Pajarinen, Gerhard Neumann, J Andrew Bagnell, Pieter Abbeel, and Jan Peters · 2018
Later among the works it cites.
Decentralization of multiagent policies by learning what to communicate
James Paulos, Steven W. Chen, Daigo Shishika, and Vijay Kumar · 2018
Later among the works it cites.
QMIX: monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Mikayel Samvelyan, Christian Schröder de Witt, Gregory Farquhar, Jakob N. Foerster, and Shimon Whiteson · 2018
Later among the works it cites.
Autonomously reusing knowledge in multiagent reinforcement learning
Felipe Leno Da Silva, Matthew E. Taylor, and Anna Helena Reali Costa · 2018
Later among the works it cites.
Multi-agent generative adversarial imitation learning
Jiaming Song, Hongyu Ren, Dorsa Sadigh, and Stefano Ermon · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Later among the works it cites.
Action branching architectures for deep reinforcement learning
Arash Tavakoli, Fabio Pardo, and Petar Kormushev · 2018
Later among the works it cites.
Value propagation for decentralized networked deep multi-agent reinforcement learning
Chao Qu, Shie Mannor, Huan Xu, Yuan Qi, Le Song, and Junwu Xiong · 2019
Closest in time.
The starcraft multi-agent challenge
Mikayel Samvelyan, Tabish Rashid, Christian Schroeder De Witt, Gregory Farquhar, Nantas Nardelli, Tim GJ Rudner, Chia-Man Hung, Philip HS Torr, Jakob Foerster, and Shimon Whiteson · 2019
Closest in time.
Multi-agent adversarial inverse reinforcement learning
Lantao Yu, Jiaming Song, and Stefano Ermon · 2019
Closest in time.
Multi-agent interactions modeling with correlated policies
Minghuan Liu, Ming Zhou, Weinan Zhang, Yuzheng Zhuang, Jun Wang, Wulong Liu, and Yong Yu · 2020
Closest in time.
A generalized training approach for multiagent learning
Paul Muller, Shayegan Omidshafiei, Mark Rowland, Karl Tuyls, Julien Perolat, Siqi Liu, Daniel Hennes, Luke Marris, Marc Lanctot, Edward Hughes, Zhe Wang, Guy Lever, Nicolas Heess, Thore Graepel, and Remi Munos · 2020
Closest in time.