Fetching the paper…
Reading the bibliography…
We explore deep reinforcement learning methods for multi-agent domains.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Feudal reinforcement learning
P. Dayan and G. E. Hinton · 1993
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
M. Tan · 1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
M. L. Littman · 1994
Earlier work this paper cites.
Learning conventions in multiagent stochastic domains using likelihood estimates
C. Boutilier · 1996
Earlier work this paper cites.
Online learning about other agents in a dynamic multiagent system
J. Hu and M. P. Wellman · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
An algorithm for distributed reinforcement learning in cooperative multi-agent systems
M. Lauer and M. Riedmiller · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
Coordination in multiagent reinforcement learning: a bayesian approach
G. Chalkiadakis and C. Boutilier · 2003
Earlier work this paper cites.
Extending q-learning to general adaptive multi-agent systems
G. Tesauro · 2004
Earlier work this paper cites.
Cooperative multi-agent learning: The state of the art
L. Panait and S. Luke · 2005
Earlier work this paper cites.
Hysteretic q-learning: an algorithm for decentralized reinforcement learning in cooperative multi-agent teams
L. Matignon, G. J. Laurent, and N. Le Fort-Piat · 2007
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
L. Busoniu, R. Babuska, and B. De Schutter · 2008
Cited alongside, same era.
Conjugate markov decision processes
P. S. Thomas and A. G. Barto · 2011
Cited alongside, same era.
Predicting pragmatic reasoning in language games
M. C. Frank and N. D. Goodman · 2012
Cited alongside, same era.
Coordinated multi-robot exploration under communication constraints using decentralized markov decision processes
L. Matignon, L. Jeanpierre, A.-I. Mouaddib, et al · 2012
Cited alongside, same era.
Independent reinforcement learners in cooperative markov games: a survey regarding coordination problems
L. Matignon, G. J. Laurent, and N. Le Fort-Piat · 2012
Cited alongside, same era.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Multi-agent cooperation and the emergence of (natural) language
A. Lazaridou, A. Peysakhovich, and M. Baroni · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Later among the works it cites.
Learning multiagent communication with backpropagation
S. Sukhbaatar, R. Fergus, et al · 2016
Later among the works it cites.
https://deepmind.com/blog/deepmind-ai-reduces-google-data-centre-cooling-bill-40/
DeepMind AI reduces google data centre cooling bill by 40 · 2017
Closest in time.
Counterfactual multi-agent policy gradients
J. Foerster, G. Farquhar, T. Afouras, N. Nardelli, and S. Whiteson · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Cited alongside, same era.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Learning to protect communications with adversarial neural cryptography
M. Abadi and D. G. Andersen · 2016
Cited alongside, same era.
Learning to communicate with deep multi-agent reinforcement learning
J. N. Foerster, Y. M. Assael, N. de Freitas, and S. Whiteson · 2016
Cited alongside, same era.
Closest in time.
Stabilising experience replay for deep multi-agent reinforcement learning
J. N. Foerster, N. Nardelli, G. Farquhar, P. H. S. Torr, P. Kohli, and S. Whiteson · 2017
Closest in time.
Cooperative multi-agent control using deep reinforcement learning
J. K. Gupta, M. Egorov, and M. Kochenderfer · 2017
Closest in time.
Multi-agent reinforcement learning in sequential social dilemmas
J. Z. Leibo, V. F. Zambaldi, M. Lanctot, J. Marecki, and T. Graepel · 2017
Closest in time.
Emergence of grounded compositional language in multi-agent populations
I. Mordatch and P. Abbeel · 2017
Closest in time.
Deep decentralized multi-task multi-agent reinforcement learning under partial observability
S. Omidshafiei, J. Pazis, C. Amato, J. P. How, and J. Vian · 2017
Closest in time.
Multiagent bidirectionally-coordinated nets for learning to play starcraft combat games
P. Peng, Q. Yuan, Y. Wen, Y. Yang, Z. Tang, H. Long, and J. Wang · 2017
Closest in time.
Intrinsic motivation and automatic curricula via asymmetric self-play
S. Sukhbaatar, I. Kostrikov, A. Szlam, and R. Fergus · 2017
Closest in time.
Multiagent cooperation and competition with deep reinforcement learning
A. Tampuu, T. Matiisen, D. Kodelja, I. Kuzovkin, K. Korjus, J. Aru, J. Aru, and R. Vicente · 2017
Closest in time.