Fetching the paper…
Reading the bibliography…
Cooperative multi-agent reinforcement learning (MARL) is making rapid progress for solving tasks in a grid world and real-world scenarios, in which agents are given different attributes and goals, resulting in different behavior through the whole multi-agent task.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L. J · 1992
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Tan, M · 1993
Earlier work this paper cites.
Dynamic noncooperative game theory
Başar, T. and Olsder, G. J · 1998
Earlier work this paper cites.
Friend-or-foe q-learning in general-sum games
Littman, M. L · 2001
Earlier work this paper cites.
Neural Network Learning - Theoretical Foundations
Anthony, M. and Bartlett, P. L · 2002
Earlier work this paper cites.
Nash q-learning for general-sum stochastic games
Hu, J. and Wellman, M. P · 2003
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., Geurts, P., and Wehenkel, L · 2005
Earlier work this paper cites.
Neural fitted Q iteration - first experiences with a data efficient neural reinforcement learning method
Riedmiller, M. A · 2005
Earlier work this paper cites.
Towards understanding linear value decomposition in cooperative multi-agent q-learning
Wang, J., Ren, Z., Han, B., Ye, J., and Zhang, C · 2006
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Munos, R. and Szepesvári, C · 2008
Earlier work this paper cites.
Error propagation for approximate policy and value iteration
Farahmand, A. M., Munos, R., and Szepesvári, C · 2010
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, S., He, K., Girshick, R., and Sun, J · 2015
Earlier work this paper cites.
Approximate modified policy iteration and its application to the game of tetris
Scherrer, B., Ghavamzadeh, M., Gabillon, V., Lesner, B., and Geist, M · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
Regularized policy iteration with nonparametric function spaces
Farahmand, A.-m., Ghavamzadeh, M., Szepesvári, C., and Mannor, S · 2016
Earlier work this paper cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Kulkarni, T. D., Narasimhan, K., Saeedi, A., and Tenenbaum, J · 2016
Earlier work this paper cites.
Analysis of classification-based policy iteration algorithms
Lazaric, A., Ghavamzadeh, M., and Munos, R · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
A concise introduction to decentralized POMDPs , volume 1
Oliehoek, F. A., Amato, C., et al · 2016
Earlier work this paper cites.
You only look once: Unified, real-time object detection
Redmon, J., Divvala, S., Girshick, R., and Farhadi, A · 2016
Earlier work this paper cites.
Learning multiagent communication with backpropagation
Sukhbaatar, S., Szlam, A., and Fergus, R · 2016
Earlier work this paper cites.
Coordinated multi-agent imitation learning
Le, H. M., Yue, Y., Carr, P., and Lucey, P · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R., Wu, Y., Tamar, A., Harb, J., Abbeel, P., and Mordatch, I · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al · 2017
Cited alongside, same era.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Anderson, P., Wu, Q., Teney, D., Bruce, J., Johnson, M., Sünderhauf, N., Reid, I. D., Gould, S., and van den Hengel, A · 2018
The hanabi challenge: A new frontier for ai research
Bard, N., Foerster, J. N., Chandar, S., Burch, N., Lanctot, M., Song, H. F., Parisotto, E., Dumoulin, V., Moitra, S., Hughes, E., et al · 2020
Later among the works it cites.
A theoretical analysis of deep q-learning
Fan, J., Wang, Z., Xie, Y., and Yang, Z · 2020
Later among the works it cites.
Google research football: A novel reinforcement learning environment
Kurach, K., Raichuk, A., Stanczyk, P., Zajac, M., Bachem, O., Espeholt, L., Riquelme, C., Vincent, D., Michalski, M., Bousquet, O., and Gelly, S · 2020
Later among the works it cites.
Emergent multi-agent communication in the deep learning era
Lazaridou, A. and Baroni, M · 2020
Later among the works it cites.
Suphx: Mastering mahjong with deep reinforcement learning
Li, J., Koyamada, S., Ye, Q., Liu, G., Wang, C., Yang, R., Zhao, L., Qin, T., Liu, T.-Y., and Hon, H.-W · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Counterfactual multi-agent policy gradients
Foerster, J. N., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D · 2018
Cited alongside, same era.
Learning attentional communication for multi-agent cooperation
Jiang, J. and Lu, Z · 2018
Cited alongside, same era.
QMIX: monotonic value function factorisation for deep multi-agent reinforcement learning
Rashid, T., Samvelyan, M., de Witt, C. S., Farquhar, G., Foerster, J. N., and Whiteson, S · 2018
Cited alongside, same era.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W. M., Zambaldi, V. F., Jaderberg, M., Lanctot, M., Sonnerat, N., Leibo, J. Z., Tuyls, K., and Graepel, T · 2018
Cited alongside, same era.
Mean field multi-agent reinforcement learning
Yang, Y., Luo, R., Li, M., Zhou, M., Zhang, W., and Wang, J · 2018
Cited alongside, same era.
Nonparametric regression using deep neural networks with relu activation function
Schmidt-Hieber, J · 2020
Later among the works it cites.
Revisiting parameter sharing in multi-agent deep reinforcement learning
Terry, J. K., Grammel, N., Hari, A., Santos, L., and Black, B · 2020
Later among the works it cites.
Mastering complex control in moba games with deep reinforcement learning
Ye, D., Liu, Z., Sun, M., Shi, B., Zhao, P., Wu, H., Yu, H., Yang, S., Wu, X., Guo, Q., et al · 2020
Later among the works it cites.
Smarts: Scalable multi-agent reinforcement learning training school for autonomous driving
Zhou, M., Luo, J., Villella, J., Yang, Y., Rusu, D., Miao, J., Zhang, W., Alban, M., Fadakar, I., Chen, Z., et al · 2020
Later among the works it cites.
Scaling multi-agent reinforcement learning with selective parameter sharing
Christianos, F., Papoudakis, G., Rahman, M. A., and Albrecht, S. V · 2021
Later among the works it cites.
Updet: Universal multi-agent RL via policy decoupling with transformers
Hu, S., Zhu, F., Chang, X., and Liang, X · 2021
Later among the works it cites.
Learning in nonzero-sum stochastic games with potentials
Mguni, D. H., Wu, Y., Du, Y., Yang, Y., Wang, Z., Li, M., Wen, Y., Jennings, J., and Wang, J · 2021
Later among the works it cites.
Benchmarking multi-agent deep reinforcement learning algorithms in cooperative tasks
Papoudakis, G., Christianos, F., Schäfer, L., and Albrecht, S. V · 2021
Later among the works it cites.
The neural MMO platform for massively multiagent research
Suarez, J., Du, Y., Zhu, C., Mordatch, I., and Isola, P · 2021
Later among the works it cites.
Game plan: What ai can do for football, and what football can do for ai
Tuyls, K., Omidshafiei, S., Muller, P., Wang, Z., Connor, J., Hennes, D., Graham, I., Spearman, W., Waskett, T., Steel, D., et al · 2021
Later among the works it cites.
RODE: learning roles to decompose multi-agent tasks
Wang, T., Gupta, T., Mahajan, A., Peng, B., Whiteson, S., and Zhang, C · 2021
Later among the works it cites.
The surprising effectiveness of mappo in cooperative, multi-agent games
Yu, C., Velu, A., Vinitsky, E., Wang, Y., Bayen, A., and Wu, Y · 2021
Later among the works it cites.
Douzero: Mastering doudizhu with self-play deep reinforcement learning
Zha, D., Xie, J., Ma, W., Zhang, S., Lian, X., Hu, X., and Liu, J · 2021
Later among the works it cites.
Finite-sample analysis for decentralized batch multi-agent reinforcement learning with networked agents
Zhang, K., Yang, Z., Liu, H., Zhang, T., and Basar, T · 2021
Later among the works it cites.
Towards distraction-robust active visual tracking
Zhong, F., Sun, P., Luo, W., Yan, T., and Wang, Y · 2021
Later among the works it cites.
Main: A multi-agent indoor navigation benchmark for cooperative learning
Zhu, F., Hu, S., Zhang, Y., Hong, H., Zhu, Y., Chang, X., and Liang, X · 2021
Later among the works it cites.