Fetching the paper…
Reading the bibliography…
A variety of cooperative multi-agent control problems require agents to achieve individual goals while contributing to collective success.
Zhang, Z., Yang, J., and Zha, H. (2019) · 1909
Earlier work this paper cites.
Efficient exploration in reinforcement learning
Thrun, S. B. (1992) · 1992
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Tan, M. (1993) · 1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Littman, M. L. (1994) · 1994
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S. (1999) · 1999
Earlier work this paper cites.
Multiagent systems: A survey from a machine learning perspective
Stone, P. and Veloso, M. (2000) · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
The communicative multiagent team decision problem: Analyzing teamwork theories and models
Pynadath, D. V. and Tambe, M. (2002) · 2002
Earlier work this paper cites.
Nash q-learning for general-sum stochastic games
Hu, J. and Wellman, M. P. (2003) · 2003
Earlier work this paper cites.
Multi-agent reinforcement learning: a critical survey
Shoham, Y., Powers, R., and Grenager, T. (2003) · 2003
Earlier work this paper cites.
All learning is local: Multi-agent learning in global reward games
Chang, Y.-H., Ho, T., and Kaelbling, L. P. (2004) · 2004
Earlier work this paper cites.
Cooperative multi-agent learning: The state of the art
Panait, L. and Luke, S. (2005) · 2005
Earlier work this paper cites.
Optimal and approximate q-value functions for decentralized pomdps
Oliehoek, F. A., Spaan, M. T., and Vlassis, N. (2008) · 2008
Earlier work this paper cites.
Curriculum learning
Bengio, Y., Louradour, J., Collobert, R., and Weston, J. (2009) · 2009
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
Taylor, M. E. and Stone, P. (2009) · 2009
Earlier work this paper cites.
Multi-agent path planning with multiple tasks and distance constraints
Bhattacharya, S., Likhachev, M., and Kumar, V. (2010) · 2010
Earlier work this paper cites.
A survey on transfer learning
Pan, S. J. and Yang, Q. (2010) · 2010
Earlier work this paper cites.
An overview of recent progress in the study of distributed multi-agent coordination
Cao, Y., Yu, W., Ren, W., and Chen, G. (2013) · 2013
Earlier work this paper cites.
The psychology of social dilemmas: A review
Van Lange, P. A., Joireman, J., Parks, C. D., and Van Dijk, E. (2013) · 2013
Cited alongside, same era.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. (2014) · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D. (2015) · 2015
Cited alongside, same era.
How other-regarding preferences can promote cooperation in non-zero-sum grid games
Austerweil, J. L., Brawner, S., Greenwald, A., Hilliard, E., Ho, M., Littman, M. L., MacGlashan, J., and Trimbach, C. (2016) · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Microscopic traffic simulation using SUMO
Lopez, P. A., Behrisch, M., Bieker-Walz, L., Erdmann, J., Flötteröd, Y.-P., Hilbrich, R., Lücken, L., Rummel, J., Wagner, P., and Wießner, E. (2018) · 2018
Closest in time.
Emergence of grounded compositional language in multi-agent populations
Mordatch, I. and Abbeel, P. (2018) · 2018
Closest in time.
Credit assignment for collective multiagent rl with global rewards
Nguyen, D. T., Kumar, A., and Lau, H. C. (2018) · 2018
Closest in time.
QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning
Rashid, T., Samvelyan, M., Schroeder, C., Farquhar, G., Foerster, J., and Whiteson, S. (2018) · 2018
Closest in time.
Actor-critic policy optimization in partially observable multiagent environments
Srinivasan, S., Lanctot, M., Zambaldi, V., Pérolat, J., Tuyls, K., Munos, R., and Bowling, M. (2018) · 2018
Closest in time.
Variance reduction for policy gradient with action-dependent factorized baselines
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2016) · 2016
Cited alongside, same era.
Rusu, A. A., Rabinowitz, N. C., Desjardins, G., Soyer, H., Kirkpatrick, J., Kavukcuoglu, K., Pascanu, R., and Hadsell, R. (2016) · 2016
Cited alongside, same era.
Learning multiagent communication with backpropagation
Sukhbaatar, S., Fergus, R., et al · 2016
Cited alongside, same era.
Cooperative multi-agent control using deep reinforcement learning
Gupta, J. K., Egorov, M., and Kochenderfer, M. (2017) · 2017
Cited alongside, same era.
Imitating driver behavior with generative adversarial networks
Kuefler, A., Morton, J., Wheeler, T., and Kochenderfer, M. (2017) · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R., Wu, Y., Tamar, A., Harb, J., Abbeel, O. P., and Mordatch, I. (2017) · 2017
Cited alongside, same era.
Deep decentralized multi-task multi-agent reinforcement learning under partial observability
Omidshafiei, S., Pazis, J., Amato, C., How, J. P., and Vian, J. (2017) · 2017
Cited alongside, same era.
Wu, C., Rajeswaran, A., Duan, Y., Kumar, V., Bayen, A. M., Kakade, S., Mordatch, I., and Abbeel, P. (2018) · 2018
Closest in time.
Fully decentralized multi-agent reinforcement learning with networked agents
Zhang, K., Yang, Z., Liu, H., Zhang, T., and Basar, T. (2018) · 2018
Closest in time.
Starcraft ii
Blizzard Entertainment (2019) · 2019
Closest in time.
A structured prediction approach for generalization in cooperative multi-agent reinforcement learning
Carion, N., Synnaeve, G., Lazaric, A., and Usunier, N. (2019) · 2019
Closest in time.
Human-level performance in 3d multiplayer games with population-based reinforcement learning
Jaderberg, M., Czarnecki, W. M., Dunning, I., Marris, L., Lever, G., Castaneda, A. G., Beattie, C., Rabinowitz, N. C., Morcos, A. S., Ruderman, A., et al · 2019
Closest in time.
Robust multi-agent reinforcement learning via minimax deep deterministic policy gradient
Li, S., Wu, Y., Cui, X., Dong, H., Fang, F., and Russell, S. (2019) · 2019
Closest in time.
Emergent coordination through competition
Liu, S., Lever, G., Merel, J., Tunyasuvunakool, S., Heess, N., and Graepel, T. (2019) · 2019
Closest in time.
M3rl: Mind-aware multi-agent management reinforcement learning
Shu, T. and Tian, Y. (2019) · 2019
Closest in time.
Learning when to communicate at scale in multiagent cooperative and competitive tasks
Singh, A., Jain, T., and Sukhbaatar, S. (2019) · 2019
Closest in time.
Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning
Son, K., Kim, D., Kang, W. J., Hostallero, D., and Yi, Y. (2019) · 2019
Closest in time.
Navigating occluded intersections with autonomous vehicles using deep reinforcement learning
Isele, D., Rahimi, R., Cosgun, A., Subramanian, K., and Fujimura, K. (2018) · 2039
Closest in time.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W. M., Zambaldi, V., Jaderberg, M., Lanctot, M., Sonnerat, N., Leibo, J. Z., Tuyls, K., et al · 2087
Closest in time.