Fetching the paper…
Reading the bibliography…
Multi-agent deep reinforcement learning (MARL) suffers from a lack of commonly-used evaluation tasks and criteria, making comparisons between approaches difficult.
Non-cooperative games
John Nash · 1951
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S. Sutton · 1988
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Long-Ji Lin · 1992
Earlier work this paper cites.
Q-learning
Christopher J. C. H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Ming Tan · 1993
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Caroline Claus and Craig Boutilier · 1998
Earlier work this paper cites.
Analyzing and visualizing multiagent rewards in dynamic and stochastic domains
Adrian K Agogino and Kagan Tumer · 2008
Earlier work this paper cites.
Double q-learning
Hado Hasselt · 2010
Earlier work this paper cites.
Comparative evaluation of MAL algorithms in a diverse set of ad hoc team problems
Stefano V. Albrecht and Subramanian Ramamoorthy · 2012
Earlier work this paper cites.
A game-theoretic model and best-response learning method for ad hoc coordination in multiagent systems
Stefano V. Albrecht and Subramanian Ramamoorthy · 2013
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Charles Beattie, Joel Z. Leibo, Denis Teplyashin, Tom Ward, Marcus Wainwright, Heinrich Küttler, Andrew Lefrancq, Simon Green, Víctor Valdés, Amir Sadik, et al · 2016
Earlier work this paper cites.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Earlier work this paper cites.
Learning to communicate with deep multi-agent reinforcement learning
Jakob N. Foerster, Ioannis Alexandros Assael, Nando De Freitas, and Shimon Whiteson · 2016
Earlier work this paper cites.
Opponent modeling in deep reinforcement learning
He He, Jordan Boyd-Graber, Kevin Kwok, and Hal Daumé III · 2016
Earlier work this paper cites.
The malmo platform for artificial intelligence experimentation
Matthew Johnson, Katja Hofmann, Tim Hutton, and David Bignell · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Learning multiagent communication with backpropagation
Sainbayar Sukhbaatar, Rob Fergus, et al · 2016
Cited alongside, same era.
Reasoning about hypothetical agent behaviours and their parameters
Stefano V. Albrecht and Peter Stone · 2017
Cited alongside, same era.
OpenAI baselines, 2017
Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, Yuhuai Wu, and Peter Zhokhov · 2017
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2017
LIIR: Learning individual intrinsic reward in multi-agent reinforcement learning
Yali Du, Lei Han, Meng Fang, Ji Liu, Tianhong Dai, and Dacheng Tao · 2019
Later among the works it cites.
Implementation matters in deep RL: A case study on PPO and TRPO
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, and Aleksander Madry · 2019
Later among the works it cites.
A survey and critique of multiagent deep reinforcement learning
Pablo Hernandez-Leal, Bilal Kartal, and Matthew E. Taylor · 2019
Later among the works it cites.
Actor-attention-critic for multi-agent reinforcement learning
Shariq Iqbal and Fei Sha · 2019
Later among the works it cites.
Social influence as intrinsic motivation for multi-agent deep reinforcement learning
Natasha Jaques, Angeliki Lazaridou, Edward Hughes, Caglar Gulcehre, Pedro A. Ortega, DJ Strouse, Joel Z. Leibo, and Nando De Freitas · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch · 2017
Cited alongside, same era.
The concrete distribution: A continuous relaxation of discrete random variables
Chris J. Maddison, Andriy Mnih, and Yee Whye Teh · 2017
Cited alongside, same era.
Emergence of grounded compositional language in multi-agent populations
Igor Mordatch and Pieter Abbeel · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
StarCraft II: A new challenge for reinforcement learning
Oriol Vinyals, Timo Ewalds, Sergey Bartunov, Petko Georgiev, Alexander Sasha Vezhnevets, Michelle Yeo, Alireza Makhzani, Heinrich Küttler, John Agapiou, Julian Schrittwieser, et al · 2017
Cited alongside, same era.
Counterfactual multi-agent policy gradients
Jakob N. Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2018
Cited alongside, same era.
Georgios Papoudakis, Filippos Christianos, Arrasy Rahman, and Stefano V. Albrecht · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Later among the works it cites.
The StarCraft multi-agent challenge
Mikayel Samvelyan, Tabish Rashid, Christian Schroeder de Witt, Gregory Farquhar, Nantas Nardelli, Tim GJ Rudner, Chia-Man Hung, Philip HS Torr, Jakob Foerster, and Shimon Whiteson · 2019
Later among the works it cites.
Benchmarking model-based reinforcement learning
Tingwu Wang, Xuchan Bao, Ignasi Clavera, Jerrick Hoang, Yeming Wen, Eric Langlois, Shunshi Zhang, Guodong Zhang, Pieter Abbeel, and Jimmy Ba · 2019
Later among the works it cites.
The hanabi challenge: A new frontier for ai research
Nolan Bard, Jakob N. Foerster, Sarath Chandar, Neil Burch, Marc Lanctot, H. Francis Song, Emilio Parisotto, Vincent Dumoulin, Subhodeep Moitra, Edward Hughes, et al · 2020
Closest in time.
Shared experience actor-critic for multi-agent reinforcement learning
Filippos Christianos, Lukas Schäfer, and Stefano V. Albrecht · 2020
Closest in time.
Multi-agent actor-critic with hierarchical graph attention network
Heechang Ryu, Hayong Shin, and Jinkyoo Park · 2020
Closest in time.
What matters in on-policy reinforcement learning? a large-scale empirical study
Marcin Andrychowicz, Anton Raichuk, Piotr Stańczyk, Manu Orsini, Sertan Girgin, Raphael Marinier, Léonard Hussenot, Matthieu Geist, Olivier Pietquin, Marcin Michalski, et al · 2021
Closest in time.
Scaling multi-agent reinforcement learning with selective parameter sharing
Filippos Christianos, Georgios Papoudakis, Arrasy Rahman, and Stefano V. Albrecht · 2021
Closest in time.
Local information agent modelling in partially-observable environments
Georgios Papoudakis, Filippos Christianos, and Stefano V. Albrecht · 2021
Closest in time.
Towards open ad hoc teamwork using graph-based policy learning
Arrasy Rahman, Niklas Höpner, Filippos Christianos, and Stefano V. Albrecht · 2021
Closest in time.
The surprising effectiveness of mappo in cooperative, multi-agent games
Chao Yu, Akash Velu, Eugene Vinitsky, Yu Wang, Alexandre Bayen, and Yi Wu · 2021
Closest in time.