Fetching the paper…
Reading the bibliography…
Multi-agent RL is rendered difficult due to the non-stationary nature of environment perceived by individual agents.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 1935
Earlier work this paper cites.
Robocup: The robot world cup initiative
Hiroaki Kitano, Minoru Asada, Yasuo Kuniyoshi, Itsuki Noda, and Eiichi Osawa · 1997
Earlier work this paper cites.
"other-play" for zero-shot coordination
Hengyuan Hu, Adam Lerer, Alex Peysakhovich, and Jakob N. Foerster · 2003
Earlier work this paper cites.
Learning to cooperate: Emergent communication in multi-agent navigation
Ivana Kajic, Eser Aygün, and Doina Precup · 2004
Earlier work this paper cites.
Half field offense in RoboCup soccer: A multiagent reinforcement learning case study
Shivaram Kalyanakrishnan, Yaxin Liu, and Peter Stone · 2007
Earlier work this paper cites.
Mastering Atari with Discrete World Models
Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2010
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller · 2013
Earlier work this paper cites.
Learning to play guess who? and inventing a grounded language as a consequence
Emilio Jorge, Mikael Kågebäck, and Emil Gustavsson · 2016
Earlier work this paper cites.
Learning to communicate with deep multi-agent reinforcement learning
Jakob Foerster, Ioannis Alexandros Assael, Nando de Freitas, and Shimon Whiteson · 2016
Cited alongside, same era.
Emergence of language with multi-agent games: Learning to communicate with sequences of symbols
Serhii Havrylov and Ivan Titov · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, YI WU, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax, 2017
Eric Jang, Shixiang Gu, and Ben Poole · 2017
Cited alongside, same era.
Emergence of linguistic communication from referential games with symbolic and pixel input
Angeliki Lazaridou, Karl Moritz Hermann, Karl Tuyls, and Stephen Clark · 2018
Later among the works it cites.
Emergent communication in a multi-modal, multi-step referential game
Katrina Evtimova, Andrew Drozdov, Douwe Kiela, and Kyunghyun Cho · 2018
Later among the works it cites.
Learning latent dynamics for planning from pixels
Danijar Hafner, Timothy P. Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson · 2018
Later among the works it cites.
Deep reinforcement learning for sequence-to-sequence models
Yaser Keneshloo, Tian Shi, Naren Ramakrishnan, and Chandan K. Reddy · 2019
Later among the works it cites.
Emergent coordination through competition
Siqi Liu, Guy Lever, Nicholas Heess, Josh Merel, Saran Tunyasuvunakool, and Thore Graepel · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kris Cao, Angeliki Lazaridou, Marc Lanctot, Joel Z Leibo, Karl Tuyls, and Stephen Clark · 2018
Cited alongside, same era.
Emergence of grounded compositional language in multi-agent populations
Igor Mordatch and Pieter Abbeel · 2018
Cited alongside, same era.
Dream to control: Learning behaviors by latent imagination
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi
Cited in the paper.
Yoshua Bengio, Nicholas Léonard, and Aaron C. Courville · 2019
Later among the works it cites.
Mastering Atari, Go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, Timothy Lillicrap, and David Silver · 2020
Later among the works it cites.