Fetching the paper…
Reading the bibliography…
Reinforcement learning in multi-agent scenarios is important for real-world applications but presents challenges beyond those seen in single-agent settings.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Multi-agent reinforcement learning: independent versus cooperative agents
Tan, M · 1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Littman, M. L · 1994
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
Hierarchical reinforcement learning in communication-mediated multiagent coordination
Fischer, F., Rovatsos, M., and Weiss, G · 2004
Earlier work this paper cites.
Multi-agent reinforcement learning: An overview
Buşoniu, L., Babuška, R., and De Schutter, B · 2010
Earlier work this paper cites.
Graves, A., Wayne, G., and Danihelka, I · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Recurrent models of visual attention
Mnih, V., Heess, N., Graves, A., et al · 2014
Earlier work this paper cites.
Multiple object recognition with visual attention
Ba, J., Mnih, V., and Kavukcuoglu, K · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
Learning to communicate with deep multi-agent reinforcement learning
Foerster, J., Assael, I. A., de Freitas, N., and Whiteson, S · 2016
Cited alongside, same era.
Opponent modeling in deep reinforcement learning
He, H., Boyd-Graber, J., Kwok, K., and Daumé III, H · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Cited alongside, same era.
Control of memory, active perception, and action in minecraft
Oh, J., Chockalingam, V., Lee, H., et al · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
A structured self-attentive sentence embedding
Lin, Z., Feng, M., Santos, C. N. d., Yu, M., Xiang, B., Zhou, B., and Bengio, Y · 2017
Later among the works it cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R., Wu, Y., Tamar, A., Harb, J., Abbeel, O. P., and Mordatch, I · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
Multiagent cooperation and competition with deep reinforcement learning
Tampuu, A., Matiisen, T., Kodelja, D., Kuzovkin, I., Korjus, K., Aru, J., Aru, J., and Vicente, R · 2017
Later among the works it cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning multiagent communication with backpropagation
Sukhbaatar, S., Fergus, R., et al · 2016
Cited alongside, same era.
Multi-focus attention network for efficient deep reinforcement learning
Choi, J., Lee, B.-J., and Zhang, B.-T · 2017
Cited alongside, same era.
Stabilising experience replay for deep multi-agent reinforcement learning
Foerster, J., Nardelli, N., Farquhar, G., Afouras, T., Torr, P. H. S., Kohli, P., and Whiteson, S · 2017
Cited alongside, same era.
Cooperative multi-agent control using deep reinforcement learning
Gupta, J. K., Egorov, M., and Kochenderfer, M · 2017
Cited alongside, same era.
Emergence of locomotion behaviours in rich environments
Heess, N., Sriram, S., Lemmon, J., Merel, J., Wayne, G., Tassa, Y., Erez, T., Wang, Z., Eslami, A., Riedmiller, M., et al · 2017
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Jang, E., Gu, S., and Poole, B · 2017
Cited alongside, same era.
Foerster, J., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S · 2018
Closest in time.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Closest in time.
Learning attentional communication for multi-agent cooperation
Jiang, J. and Lu, Z · 2018
Closest in time.
Emergence of grounded compositional language in multi-agent populations
Mordatch, I. and Abbeel, P · 2018
Closest in time.
QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning
Rashid, T., Samvelyan, M., Schroeder, C., Farquhar, G., Foerster, J., and Whiteson, S · 2018
Closest in time.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W. M., Zambaldi, V., Jaderberg, M., Lanctot, M., Sonnerat, N., Leibo, J. Z., Tuyls, K., and Graepel, T · 2018
Closest in time.
Wei, E., Wicke, D., Freelan, D., and Luke, S · 2018
Closest in time.