Fetching the paper…
Reading the bibliography…
We investigate how reinforcement learning agents can learn to cooperate.
Theory of the firm: Managerial behavior, agency costs and ownership structure
Jensen, M. C. and Meckling, W. H · 1976
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Feudal reinforcement learning
Dayan, P. and Hinton, G. E · 1993
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Tan, M · 1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Littman, M. L · 1994
Earlier work this paper cites.
Planning, learning and coordination in multiagent decision processes
Boutilier, C · 1996
Earlier work this paper cites.
Multiagent reinforcement learning: theoretical framework and an algorithm
Hu, J., Wellman, M. P., et al · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Dietterich, T. G · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
Hierarchical multi-agent reinforcement learning
Makar, R., Mahadevan, S., and Ghavamzadeh, M · 2001
Earlier work this paper cites.
Optimal payoff functions for members of collectives
Wolpert, D. H. and Tumer, K · 2002
Cited alongside, same era.
Recent advances in hierarchical reinforcement learning
Barto, A. G. and Mahadevan, S · 2003
Cited alongside, same era.
All learning is local: Multi-agent learning in global reward games
Chang, Y.-H., Ho, T., and Kaelbling, L. P · 2004
Cited alongside, same era.
Cooperative multi-agent learning: The state of the art
Panait, L. and Luke, S · 2005
Cited alongside, same era.
A concise introduction to multiagent systems and distributed artificial intelligence
Vlassis, N · 2007
Cited alongside, same era.
A comprehensive survey of multiagent reinforcement learning
Busoniu, L., Babuska, R., and De Schutter, B · 2008
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Later among the works it cites.
Learning to communicate with deep multi-agent reinforcement learning
Foerster, J., Assael, I. A., de Freitas, N., and Whiteson, S · 2016
Later among the works it cites.
Learning multiagent communication with backpropagation
Sukhbaatar, S., Fergus, R., et al · 2016
Later among the works it cites.
Cooperative multi-agent control using deep reinforcement learning
Gupta, J. K., Egorov, M., and Kochenderfer, M · 2017
Later among the works it cites.
Revisiting the master-slave architecture in multi-agent deep reinforcement learning
Kong, X., Xin, B., Liu, F., and Wang, Y · 2017
Later among the works it cites.
Federated control with hierarchical multi-agent deep reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The theory of incentives: the principal-agent model
Laffont, J.-J. and Martimort, D · 2009
Cited alongside, same era.
Independent reinforcement learners in cooperative markov games: a survey regarding coordination problems
Matignon, L., Laurent, G. J., and Le Fort-Piat, N · 2012
Cited alongside, same era.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Kumar, S., Shah, P., Hakkani-Tur, D., and Heck, L · 2017
Later among the works it cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R., Wu, Y., Tamar, A., Harb, J., Abbeel, O. P., and Mordatch, I · 2017
Later among the works it cites.
Feudal networks for hierarchical reinforcement learning
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K · 2017
Later among the works it cites.
Counterfactual multi-agent policy gradients
Foerster, J. N., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S · 2018
Later among the works it cites.
Emergence of grounded compositional language in multi-agent populations
Mordatch, I. and Abbeel, P · 2018
Later among the works it cites.
Hierarchical deep multiagent reinforcement learning
Tang, H., Hao, J., Lv, T., Chen, Y., Zhang, Z., Jia, H., Ren, C., Zheng, Y., Fan, C., and Wang, L · 2018
Later among the works it cites.