Fetching the paper…
Reading the bibliography…
Reward decomposition is a critical problem in centralized training with decentralized execution~(CTDE) paradigm for multi-agent reinforcement learning.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Tan, M · 1993
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Claus, C. and Boutilier, C · 1998
Earlier work this paper cites.
Online meta-critic learning for off-policy actor-critic methods
Zhou, W., Li, Y., Yang, Y., Wang, H., and Hospedales, T. M · 2003
Earlier work this paper cites.
Learning implicit credit assignment for multi-agent actor-critic
Zhou, M., Liu, Z., Sui, P., Li, Y., and Chung, Y. Y · 2007
Earlier work this paper cites.
Visualizing data using t-sne
Maaten, L. v. d. and Hinton, G · 2008
Earlier work this paper cites.
Qplex: Duplex dueling multi-agent q-learning
Wang, J., Ren, Z., Liu, T., Yu, Y., and Zhang, C · 2008
Earlier work this paper cites.
Rode: Learning roles to decompose multi-agent tasks
Wang, T., Gupta, T., Mahajan, A., Peng, B., Whiteson, S., and Zhang, C · 2010
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Learning to communicate with deep multi-agent reinforcement learning
Foerster, J., Assael, I. A., De Freitas, N., and Whiteson, S · 2016
Earlier work this paper cites.
A concise introduction to decentralized POMDPs , volume 1
Oliehoek, F. A., Amato, C., et al · 2016
Earlier work this paper cites.
Learning multiagent communication with backpropagation
Sukhbaatar, S., Fergus, R., et al · 2016
Earlier work this paper cites.
Stabilising experience replay for deep multi-agent reinforcement learning
Foerster, J., Nardelli, N., Farquhar, G., Afouras, T., Torr, P. H., Kohli, P., and Whiteson, S · 2017
Cited alongside, same era.
Cooperative multi-agent control using deep reinforcement learning
Gupta, J. K., Egorov, M., and Kochenderfer, M · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R., Wu, Y. I., Tamar, A., Harb, J., Abbeel, O. P., and Mordatch, I · 2017
Cited alongside, same era.
Multiagent cooperation and competition with deep reinforcement learning
Tampuu, A., Matiisen, T., Kodelja, D., Kuzovkin, I., Korjus, K., Aru, J., Aru, J., and Vicente, R · 2017
Cited alongside, same era.
Counterfactual multi-agent policy gradients
Foerster, J., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S · 2018
Cited alongside, same era.
Human-level performance in 3d multiplayer games with population-based reinforcement learning
Jaderberg, M., Czarnecki, W. M., Dunning, I., Marris, L., Lever, G., Castaneda, A. G., Beattie, C., Rabinowitz, N. C., Morcos, A. S., Ruderman, A., et al · 2019
Later among the works it cites.
Social influence as intrinsic motivation for multi-agent deep reinforcement learning
Jaques, N., Lazaridou, A., Hughes, E., Gulcehre, C., Ortega, P., Strouse, D., Leibo, J. Z., and De Freitas, N · 2019
Later among the works it cites.
Maven: Multi-agent variational exploration
Mahajan, A., Rashid, T., Samvelyan, M., and Whiteson, S · 2019
Later among the works it cites.
The starcraft multi-agent challenge
Samvelyan, M., Rashid, T., de Witt, C. S., Farquhar, G., Nardelli, N., Rudner, T. G., Hung, C.-M., Torr, P. H., Foerster, J., and Whiteson, S · 2019
Later among the works it cites.
Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rashid, T., Samvelyan, M., De Witt, C. S., Farquhar, G., Foerster, J., and Whiteson, S · 2018
Cited alongside, same era.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W. M., Zambaldi, V. F., Jaderberg, M., Lanctot, M., Sonnerat, N., Leibo, J. Z., Tuyls, K., et al · 2018
Cited alongside, same era.
Young, K., Wang, B., and Taylor, M. E · 2018
Cited alongside, same era.
On learning intrinsic rewards for policy gradient methods
Zheng, Z., Oh, J., and Singh, S · 2018
Cited alongside, same era.
Emergent tool use from multi-agent autocurricula
Baker, B., Kanitscheider, I., Markov, T., Wu, Y., Powell, G., McGrew, B., and Mordatch, I · 2019
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Dębiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al · 2019
Cited alongside, same era.
Exploration with unreliable intrinsic reward in multi-agent reinforcement learning
Böhmer, W., Rashid, T., and Whiteson, S · 2019
Cited alongside, same era.
Son, K., Kim, D., Kang, W. J., Hostallero, D. E., and Yi, Y · 2019
Later among the works it cites.
Discovery of useful questions as auxiliary tasks
Veeriah, V., Hessel, M., Xu, Z., Rajendran, J., Lewis, R. L., Oh, J., van Hasselt, H. P., Silver, D., and Singh, S · 2019
Later among the works it cites.
Learning individually inferred communication for multi-agent cooperation
Ding, Z., Huang, T., and Lu, Z · 2020
Later among the works it cites.
Weighted qmix: Expanding monotonic value function factorisation
Rashid, T., Farquhar, G., Peng, B., and Whiteson, S · 2020
Later among the works it cites.
Joint policy search for multi-agent collaboration with imperfect information
Tian, Y., Gong, Q., and Jiang, T · 2020
Later among the works it cites.
Smix ( λ \lambda ): Enhancing centralized value functions for cooperative multi-agent reinforcement learning
Wen, C., Yao, X., Wang, Y., and Tan, X · 2020
Later among the works it cites.
Qatten: A general framework for cooperative multiagent reinforcement learning
Yang, Y., Hao, J., Liao, B., Shao, K., Chen, G., Liu, W., and Tang, H · 2020
Later among the works it cites.
Towards playing full moba games with deep reinforcement learning
Ye, D., Chen, G., Zhang, W., Chen, S., Yuan, B., Liu, B., Chen, J., Liu, Z., Qiu, F., Yu, H., et al · 2020
Later among the works it cites.