Fetching the paper…
Reading the bibliography…
Existing evaluation suites for multi-agent reinforcement learning (MARL) do not assess generalization to novel situations as their primary objective (unlike supervised-learning benchmarks).
Leibo, J. Z., Hughes, E., Lanctot, M., and Graepel, T · 1903
Earlier work this paper cites.
Stochastic Games
Shapley, L. S · 1953
Earlier work this paper cites.
Games with incomplete information played by “bayesian” players, i–iii part i. the basic model
Harsanyi, J. C · 1967
Earlier work this paper cites.
The Evolution of Cooperation
Axelrod, R · 1984
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Littman, M. L · 1994
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Claus, C. and Boutilier, C · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
On six advances in cooperation theory
Axelrod, R · 2000
Earlier work this paper cites.
Mm algorithms for generalized bradley-terry models
Hunter, D. R. et al · 2004
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Lab experiments for the study of social-ecological systems
Janssen, M. A., Holahan, R., Lee, A., and Ostrom, E · 2010
Earlier work this paper cites.
Ad hoc autonomous agent teams: Collaboration without pre-coordination
Stone, P., Kaminka, G. A., Kraus, S., Rosenschein, J. S., et al · 2010
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, D · 2011
Earlier work this paper cites.
Aligning superintelligence with human interests: A technical research agenda
Soares, N. and Fallenstein, B · 2014
Earlier work this paper cites.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G · 2015
Earlier work this paper cites.
Research priorities for robust and beneficial artificial intelligence
Russell, S., Dewey, D., and Tegmark, M · 2015
Earlier work this paper cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Earlier work this paper cites.
Concrete problems in AI safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Multi-agent reinforcement learning as a rehearsal for decentralized planning
Kraemer, L. and Banerjee, B · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
A concise introduction to decentralized POMDPs
Oliehoek, F. A. and Amato, C · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Emergent complexity via multi-agent competition
Bansal, T., Pachocki, J., Sidor, S., Sutskever, I., and Mordatch, I · 2017
Earlier work this paper cites.
A unified game-theoretic approach to multiagent reinforcement learning
Lanctot, M., Zambaldi, V., Gruslys, A., Lazaridou, A., Tuyls, K., Pérolat, J., Silver, D., and Graepel, T · 2017
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R., Wu, Y., Tamar, A., Harb, J., Abbeel, P., and Mordatch, I · 2017
Earlier work this paper cites.
A multi-agent reinforcement learning model of common-pool resource appropriation
Perolat, J., Leibo, J. Z., Zambaldi, V., Beattie, C., Tuyls, K., and Graepel, T · 2017
Cited alongside, same era.
Prosocial learning agents solve generalized stag hunts better than selfish ones
Peysakhovich, A. and Lerer, A · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Cited alongside, same era.
Intrinsic motivation and automatic curricula via asymmetric self-play
Sukhbaatar, S., Lin, Z., Kostrikov, I., Synnaeve, G., Szlam, A., and Fergus, R · 2017
Cited alongside, same era.
Feudal networks for hierarchical reinforcement learning
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K · 2017
Cited alongside, same era.
On the measure of intelligence
Chollet, F · 2019
Later among the works it cites.
Clune, J · 2019
Later among the works it cites.
Quantifying generalization in reinforcement learning
Cobbe, K., Klimov, O., Hesse, C., Kim, T., and Schulman, J · 2019
Later among the works it cites.
Generalization of reinforcement learners with working and episodic memory
Fortunato, M., Tan, M., Faulkner, R., Hansen, S., Badia, A. P., Buttimore, G., Deck, C., Leibo, J. Z., and Blundell, C · 2019
Later among the works it cites.
Multi-task deep reinforcement learning with popart
Hessel, M., Soyer, H., Espeholt, L., Czarnecki, W., Schmitt, S., and van Hasselt, H · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Re-evaluating evaluation
Balduzzi, D., Tuyls, K., Perolat, J., and Graepel, T · 2018
Cited alongside, same era.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al · 2018
Cited alongside, same era.
Generalization and regularization in DQN
Farebrother, J., Machado, M. C., and Bowling, M · 2018
Cited alongside, same era.
Learning with opponent-learning awareness
Foerster, J., Chen, R. Y., Al-Shedivat, M., Whiteson, S., Abbeel, P., and Mordatch, I · 2018
Cited alongside, same era.
Inequity aversion improves cooperation in intertemporal social dilemmas
Hughes, E., Leibo, J. Z., Philips, M. G., Tuyls, K., Duéñez-Guzmán, E. A., Castañeda, A. G., Dunning, I., Zhu, T., McKee, K. R., Koster, R., Roff, H., and Graepel, T · 2018
Cited alongside, same era.
Maintaining cooperation in complex social dilemmas using deep reinforcement learning
Lerer, A. and Peysakhovich, A · 2018
Cited alongside, same era.
Emergent coordination through competition
Liu, S., Lever, G., Merel, J., Tunyasuvunakool, S., Heess, N., and Graepel, T · 2018
Cited alongside, same era.
Human-level performance in 3d multiplayer games with population-based reinforcement learning
Jaderberg, M., Czarnecki, W. M., Dunning, I., Marris, L., Lever, G., Castaneda, A. G., Beattie, C., Rabinowitz, N. C., Morcos, A. S., Ruderman, A., et al · 2019
Later among the works it cites.
Obstacle tower: A generalization challenge in vision, control, and planning
Juliani, A., Khalifa, A., Berges, V.-P., Harper, J., Teng, E., Henry, H., Crespi, A., Togelius, J., and Lange, D · 2019
Later among the works it cites.
Openspiel: A framework for reinforcement learning in games
Lanctot, M., Lockhart, E., Lespiau, J.-B., Zambaldi, V., Upadhyay, S., Pérolat, J., Srinivasan, S., Timbers, F., Tuyls, K., Omidshafiei, S., Hennes, D., Morrill, D., Muller, P., Ewalds, T., Faulkner, R., Kramar, J., De Vylder, B., Saeta, B., Bradbury, J., Ding, D., Borgeaud, S., Lai, M., Schrittwieser, J., Anthony, T., Hughes, E., Danihelka, I., and Ryan-Davis, J · 2019
Later among the works it cites.
On the pitfalls of measuring emergent communication
Lowe, R., Foerster, J., Boureau, Y.-L., Pineau, J., and Dauphin, Y · 2019
Later among the works it cites.
Behaviour suite for reinforcement learning
Osband, I., Doron, Y., Hessel, M., Aslanides, J., Sezener, E., Saraiva, A., McKinney, K., Lattimore, T., Szepesvari, C., Singh, S., et al · 2019
Later among the works it cites.
The multi-agent reinforcement learning in malmö (marlö) competition
Perez-Liebana, D., Hofmann, K., Mohanty, S. P., Kuno, N., Kramer, A., Devlin, S., Gaina, R. D., and Ionita, D · 2019
Later among the works it cites.
Automated curriculum generation through setter-solver interactions
Racaniere, S., Lampinen, A., Santoro, A., Reichert, D., Firoiu, V., and Lillicrap, T · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Later among the works it cites.
Wang, R., Lehman, J., Clune, J., and Stanley, K. O · 2019
Later among the works it cites.
Beattie, C., Köppe, T., Duéñez-Guzmán, E. A., and Leibo, J. Z · 2020
Later among the works it cites.
Open problems in cooperative AI
Dafoe, A., Hughes, E., Bachrach, Y., Collins, T., McKee, K. R., Leibo, J. Z., Larson, K., and Graepel, T · 2020
Later among the works it cites.
The weirdest people in the world: How the west became psychologically peculiar and particularly prosperous
Henrich, J · 2020
Later among the works it cites.
“other-play” for zero-shot coordination
Hu, H., Lerer, A., Peysakhovich, A., and Foerster, J · 2020
Later among the works it cites.
Model-free conventions in multi-agent reinforcement learning with heterogeneous preferences
Köster, R., McKee, K. R., Everett, R., Weidinger, L., Isaac, W. S., Hughes, E., Duéñez-Guzmán, E. A., Graepel, T., Botvinick, M., and Leibo, J. Z · 2020
Later among the works it cites.
Emergent multi-agent communication in the deep learning era
Lazaridou, A. and Baroni, M · 2020
Later among the works it cites.
The logic of universalization guides moral judgment
Levine, S., Kleiman-Weiner, M., Schulz, L., Tenenbaum, J., and Cushman, F · 2020
Later among the works it cites.
Social diversity and social preferences in mixed-motive reinforcement learning
McKee, K. R., Gemp, I., McWilliams, B., Duèñez-Guzmán, E. A., Hughes, E., and Leibo, J. Z · 2020
Later among the works it cites.
The unity game engine, 2020
Unity Technologies · 2020
Later among the works it cites.
Options as responses: Grounding behavioural hierarchies in multi-agent reinforcement learning
Vezhnevets, A., Wu, Y., Eckstein, M., Leblond, R., and Leibo, J. Z · 2020
Later among the works it cites.
Too many cooks: Coordinating multi-agent collaboration through inverse planning
Wang, R. E., Wu, S. A., Evans, J. A., Tenenbaum, J. B., Parkes, D. C., and Kleiman-Weiner, M · 2020
Later among the works it cites.
Towards playing full moba games with deep reinforcement learning
Ye, D., Chen, G., Zhang, W., Chen, S., Yuan, B., Liu, B., Chen, J., Liu, Z., Qiu, F., Yu, H., et al · 2020
Later among the works it cites.