Fetching the paper…
Reading the bibliography…
This paper investigates the use of intrinsic reward to guide exploration in multi-agent reinforcement learning.
The StarCraft multi-agent challenge
Samvelyan, M., Rashid, T., de Witt, C. S., Farquhar, G., Nardelli, N., Rudner, T. G. J., Hung, C., Torr, P. H. S., Foerster, J. N., and Whiteson, S · 1902
Earlier work this paper cites.
Leibo, J. Z., Hughes, E., Lanctot, M., and Graepel, T · 1903
Earlier work this paper cites.
Learning from delayed rewards
Watkins, C. J. C. H · 1989
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J · 1992
Earlier work this paper cites.
Q-learning
Watkins, C. and Dayan, P · 1992
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Tan, M · 1993
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, J., Gulcehre, C., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Deep recurrent q-learning for partially observable mdps
Hausknecht, M. J. and Stone, P · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M. G., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Earlier work this paper cites.
Learning to communicate with deep multi-agent reinforcement learning
Foerster, J., Assael, I. A., de Freitas, N., and Whiteson, S · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Hasselt, H. v., Guez, A., and Silver, D · 2016
Cited alongside, same era.
A concise introduction to decentralized POMDPs
Oliehoek, F. A. and Amato, C · 2016
Cited alongside, same era.
Generalization and exploration via randomized value functions
Osband, I., Van Roy, B., and Wen, Z · 2016
Cited alongside, same era.
Conditional image generation with PixelCNN decoders
van den Oord, A., Kalchbrenner, N., Espeholt, L., kavukcuoglu, k., Vinyals, O., and Graves, A · 2016
Cited alongside, same era.
Concrete dropout
Gal, Y., Hron, J., and Kendall, A · 2017
Cited alongside, same era.
Count-based exploration with neural density models
Ostrovski, G., Bellemare, M. G., van den Oord, A., and Munos, R · 2017
Cited alongside, same era.
Is Q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I · 2018
Later among the works it cites.
The uncertainty Bellman equation and exploration
O’Donoghue, B., Osband, I., Munos, R., and Mnih, V · 2018
Later among the works it cites.
Randomized prior functions for deep reinforcement learning
Osband, I., Aslanides, J., and Cassirer, A · 2018
Later among the works it cites.
Parameter space noise for exploration
Plappert, M., Houthooft, R., Dhariwal, P., Sidor, S., Chen, R. Y., Chen, X., Asfour, T., Abbeel, P., and Andrychowicz, M · 2018
Later among the works it cites.
QMIX: monotonic value function factorisation for deep multi-agent reinforcement learning
Rashid, T., Samvelyan, M., de Witt, C. S., Farquhar, G., Foerster, J. N., and Whiteson, S · 2018
Later among the works it cites.
Deep abstract Q-networks
Roderick, M., Grimm, C., and Tellex, S · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Cited alongside, same era.
#Exploration: A study of count-based exploration for deep reinforcement learning
Tang, H., Houthooft, R., Foote, D., Stooke, A., Xi Chen, O., Duan, Y., Schulman, J., DeTurck, F., and Abbeel, P · 2017
Cited alongside, same era.
Counterfactual multi-agent policy gradients
Foerster, J. N., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S · 2018
Cited alongside, same era.
Noisy networks for exploration
Fortunato, M., Azar, M. G., Piot, B., Menick, J., Hessel, M., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., Pietquin, O., Blundell, C., and Legg, S · 2018
Cited alongside, same era.
Intrinsic social motivation via causal influence in multi-agent RL
Jaques, N., Lazaridou, A., Hughes, E., Gülçehre, Ç., Ortega, P. A., Strouse, D., Leibo, J. Z., and de Freitas, N · 2018
Cited alongside, same era.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W. M., Zambaldi, V., Jaderberg, M., Lanctot, M., Sonnerat, N., Leibo, J. Z., Tuyls, K., and Graepel, T · 2018
Later among the works it cites.
Large-scale study of curiosity-driven learning
Burda, Y., Edwards, H., Pathak, D., Storkey, A., Darrell, T., and Efros, A. A · 2019
Closest in time.
AlphaStar: Mastering the real-time strategy game StarCraft II
Vinyals, O., Babuschkin, I., Chung, J., Mathieu, M., Jaderberg, M., Czarnecki, W., Dudzik, A., Huang, A., Georgiev, P., ichard Powell, Ewalds, T., Horgan, D., Kroiss, M., Danihelka, I., Agapiou, J., Oh, J., Dalibard, V., Choi, D., Sifre, L., Sulsky, Y., Vezhnevets, S., Molloy, J., Cai, T., Budden, D., Paine, T., Gulcehre, C., Wang, Z., Pfaff, T., Pohlen, T., Wu, Y., Yogatama, D., Cohen, J., McKinney, K., Smith, O., Schaul, T., Lillicrap, T., Apps, C., Kavukcuoglu, K., Hassabis, D., and Silver, D · 2019
Closest in time.
Structured exploration via hierarchical variational policy networks, 2018
Zheng, S. and Yue, Y · 2019
Closest in time.