Fetching the paper…
Reading the bibliography…
VDN and QMIX are two popular value-based algorithms for cooperative MARL that learn a centralized action value function as a monotonic mixing of per-agent utilities.
Leibo, J. Z., Hughes, E., Lanctot, M., and Graepel, T · 1903
Earlier work this paper cites.
Exploration with unreliable intrinsic reward in multi-agent reinforcement learning
Böhmer, W., Rashid, T., and Whiteson, S · 1906
Earlier work this paper cites.
Böhmer, W., Kurin, V., and Whiteson, S · 1910
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J · 1992
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Dayan, P · 1993
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Tan, M · 1993
Earlier work this paper cites.
Coordinated reinforcement learning
Guestrin, C., Lagoudakis, M., and Parr, R · 2002
Earlier work this paper cites.
Monotonic value function factorisation for deep multi-agent reinforcement learning
Rashid, T., Samvelyan, M., De Witt, C. S., Farquhar, G., Foerster, J., and Whiteson, S · 2003
Earlier work this paper cites.
Solving transition independent decentralized markov decision processes
Becker, R., Zilberstein, S., Lesser, V., and Goldman, C. V · 2004
Earlier work this paper cites.
Biasing coevolutionary search for optimal multiagent behaviors
Panait, L., Luke, S., and Wiegand, R. P · 2006
Earlier work this paper cites.
Weighted qmix: Expanding monotonic value function factorisation
Rashid, T., Farquhar, G., Peng, B., and Whiteson, S · 2006
Earlier work this paper cites.
Qplex: Duplex dueling multi-agent q-learning
Wang, J., Ren, Z., Liu, T., Yu, Y., and Zhang, C · 2008
Earlier work this paper cites.
Rode: Learning roles to decompose multi-agent tasks
Wang, T., Gupta, T., Mahajan, A., Peng, B., Whiteson, S., and Zhang, C · 2010
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, J., Gulcehre, C., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M. G., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Earlier work this paper cites.
Ha, D., Dai, A., and Le, Q. V · 2016
Earlier work this paper cites.
Multi-agent reinforcement learning as a rehearsal for decentralized planning
Kraemer, L. and Banerjee, B · 2016
Cited alongside, same era.
Deep successor reinforcement learning
Kulkarni, T. D., Saeedi, A., Gautam, S., and Gershman, S. J · 2016
Cited alongside, same era.
A concise introduction to decentralized POMDPs , volume 1
Oliehoek, F. A., Amato, C., et al · 2016
Cited alongside, same era.
Generalization and exploration via randomized value functions
Osband, I., Van Roy, B., and Wen, Z · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
van Hasselt, H., Guez, A., and Silver, D · 2016
Cited alongside, same era.
Intrinsic social motivation via causal influence in multi-agent RL
Jaques, N., Lazaridou, A., Hughes, E., Gülçehre, Ç., Ortega, P. A., Strouse, D., Leibo, J. Z., and de Freitas, N · 2018
Later among the works it cites.
The uncertainty Bellman equation and exploration
O’Donoghue, B., Osband, I., Munos, R., and Mnih, V · 2018
Later among the works it cites.
Deep abstract Q-networks
Roderick, M., Grimm, C., and Tellex, S · 2018
Later among the works it cites.
Wei, E., Wicke, D., Freelan, D., and Luke, S · 2018
Later among the works it cites.
Structured exploration via hierarchical variational policy networks, 2018
Zheng, S. and Yue, Y · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang, Z., Schaul, T., Hessel, M., Hasselt, H., Lanctot, M., and Freitas, N · 2016
Cited alongside, same era.
Lenient learning in independent-learner stochastic cooperative games
Wei, E. and Luke, S · 2016
Cited alongside, same era.
Successor features for transfer in reinforcement learning
Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., van Hasselt, H. P., and Silver, D · 2017
Cited alongside, same era.
Concrete dropout
Gal, Y., Hron, J., and Kendall, A · 2017
Cited alongside, same era.
Advantages and limitations of using successor features for transfer in reinforcement learning
Lehnert, L., Tellex, S., and Littman, M. L · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Cited alongside, same era.
Value-decomposition networks for cooperative multi-agent learning
Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W. M., Zambaldi, V., Jaderberg, M., Lanctot, M., Sonnerat, N., Leibo, J. Z., Tuyls, K., et al · 2017
Cited alongside, same era.
Large-scale study of curiosity-driven learning
Burda, Y., Edwards, H., Pathak, D., Storkey, A., Darrell, T., and Efros, A. A · 2019
Later among the works it cites.
Successor features based multi-agent rl for event-based decentralized mdps
Gupta, T., Kumar, A., and Paruchuri, P · 2019
Later among the works it cites.
Fast task inference with variational intrinsic successor features
Hansen, S., Dabney, W., Barreto, A., Van de Wiele, T., Warde-Farley, D., and Mnih, V · 2019
Later among the works it cites.
Truly batch apprenticeship learning with deep successor features
Lee, D., Srinivasan, S., and Doshi-Velez, F · 2019
Later among the works it cites.
Maven: Multi-agent variational exploration
Mahajan, A., Rashid, T., Samvelyan, M., and Whiteson, S · 2019
Later among the works it cites.
A review of cooperative multi-agent deep reinforcement learning
OroojlooyJadid, A. and Hajinezhad, D · 2019
Later among the works it cites.
The starcraft multi-agent challenge
Samvelyan, M., Rashid, T., de Witt, C. S., Farquhar, G., Nardelli, N., Rudner, T. G., Hung, C.-M., Torr, P. H., Foerster, J., and Whiteson, S · 2019
Later among the works it cites.
Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning
Son, K., Kim, D., Kang, W. J., Hostallero, D. E., and Yi, Y · 2019
Later among the works it cites.
Multiagent adversarial inverse reinforcement learning
Wei, E., Wicke, D., and Luke, S · 2019
Later among the works it cites.
Fast reinforcement learning with generalized policy updates
Barreto, A., Hou, S., Borsa, D., Silver, D., and Precup, D · 2020
Closest in time.
Qtran++: Improved value transformation for cooperative multi-agent reinforcement learning, 2020
Son, K., Ahn, S., Reyes, R. D., Shin, J., and Yi, Y · 2020
Closest in time.
Tesseract: Tensorised actors for multi-agent reinforcement learning
Mahajan, A., Samvelyan, M., Mao, L., Makoviychuk, V., Garg, A., Kossaifi, J., Whiteson, S., Zhu, Y., and Anandkumar, A · 2021
Closest in time.