Fetching the paper…
Reading the bibliography…
Collective intelligence is a fundamental trait shared by several species of living organisms.
Dota 2 with Large Scale Deep Reinforcement Learning
OpenAI, Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., Józefowicz, R., Gray, S., Olsson, C., Pachocki, J., Petrov, M., Pinto, H. P. d. O., Raiman, J., Salimans, T., Schlatter, J., Schneider, J., Sidor, S., Sutskever, I., Tang, J., Wolski, F., and Zhang, S · 1912
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Dayan, P · 1993
Earlier work this paper cites.
Collaborative plans for complex group action
Grosz, B. and Kraus, S · 1996
Earlier work this paper cites.
Towards flexible teamwork
Tambe, M · 1997
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Claus, C. and Boutilier, C · 1998
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P., Littman, M. L., and Cassandra, A. R · 1998
Earlier work this paper cites.
Defining and using ideal teammate and opponent agent models: A case study in robotic soccer
Stone, P., Riley, P., and Veloso, M · 2000
Earlier work this paper cites.
Of ants and men: Self-organized teams in human and insect organizations
Anderson, C. and McMillan, E · 2003
Earlier work this paper cites.
Predicting opponent actions by observation
Ledezma, A., Aler, R., Sanchis, A., and Borrajo, D · 2004
Earlier work this paper cites.
Coordination and adaptation in impromptu teams
Bowling, M. and McCracken, P · 2005
Earlier work this paper cites.
Learning and exploiting relative weaknesses of opponent agents
Markovitch, S. and Reger, R · 2005
Earlier work this paper cites.
A Comprehensive Survey of Multiagent Reinforcement Learning
Busoniu, L., Babuska, R., and De Schutter, B · 2008
Earlier work this paper cites.
Teamwork in self-organized robot colonies
Nouyan, S., Groß, R., Bonani, M., Mondada, F., and Dorigo, M · 2009
Earlier work this paper cites.
A masterpiece of evolution–oecophylla weaver ants (hymenoptera: Formicidae)
Crozier, R. H., Newey, P. S., Schluens, E. A., Robson, S. K., et al · 2010
Earlier work this paper cites.
Ad hoc autonomous agent teams: Collaboration without pre-coordination
Stone, P., Kaminka, G. A., Kraus, S., and Rosenschein, J. S · 2010
Earlier work this paper cites.
Empirical evaluation of ad hoc teamwork in the pursuit domain
Barrett, S., Stone, P., and Kraus, S · 2011
Earlier work this paper cites.
Efficient model learning for human-robot collaborative tasks. arxiv, 2014
Nikolaidis, S., Gu, K., Ramakrishnan, R., and Shah, J · 2014
Earlier work this paper cites.
Contextual markov decision processes, 2015
Hallak, A., Castro, D. D., and Mannor, S · 2015
Earlier work this paper cites.
Deep Recurrent Q-Learning for Partially Observable MDPs
Hausknecht, M. and Stone, P · 2015
Earlier work this paper cites.
Belief and truth in hypothesised behaviours
Albrecht, S. V., Crandall, J. W., and Ramamoorthy, S · 2016
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., Van Hasselt, H., and Silver, D · 2016
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks, 2016
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
A Concise Introduction to Decentralized POMDPs
Oliehoek, F. A. and Amato, C · 2016
Cited alongside, same era.
Efficiently detecting switches against non-stationary opponents
Hernandez-Leal, P., Zhan, Y., Taylor, M. E., Sucar, L. E., and De Cote, E. M · 2017
Cited alongside, same era.
DARLA: improving zero-shot transfer in reinforcement learning
Higgins, I., Pal, A., Rusu, A. A., Matthey, L., Burgess, C., Pritzel, A., Botvinick, M., Blundell, C., and Lerchner, A · 2017
Cited alongside, same era.
Is a good representation sufficient for sample efficient reinforcement learning?
Du, S. S., Kakade, S. M., Wang, R., and Yang, L. F · 2020
Later among the works it cites.
Monotonic value function factorisation for deep multi-agent reinforcement learning
Rashid, T., Samvelyan, M., de Witt, C. S., Farquhar, G., Foerster, J., and Whiteson, S · 2020
Later among the works it cites.
Multi-agent actor-critic with hierarchical graph attention network
Ryu, H., Shin, H., and Park, J · 2020
Later among the works it cites.
Bounds and dynamics for empirical game theoretic analysis
Tuyls, K., Perolat, J., Lanctot, M., Hughes, E., Everett, R., Leibo, J. Z., Szepesvári, C., and Graepel, T · 2020
Later among the works it cites.
OPtions as REsponses: Grounding behavioural hierarchies in multi-agent reinforcement learning
Vezhnevets, A., Wu, Y., Eckstein, M., Leblond, R., and Leibo, J. Z · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R., Wu, Y., Tamar, A., Harb, J., Abbeel, O. P., and Mordatch, I · 2017
Cited alongside, same era.
Symmetry learning for function approximation in reinforcement learning
Mahajan, A. and Tulabandhula, T · 2017
Cited alongside, same era.
Proximal Policy Optimization Algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Value-decomposition networks for cooperative multi-agent learning
Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W. M., Zambaldi, V., Jaderberg, M., Lanctot, M., Sonnerat, N., Leibo, J. Z., Tuyls, K., et al · 2017
Cited alongside, same era.
Transfer in deep reinforcement learning using successor features and generalised policy improvement
Barreto, A., Borsa, D., Quan, J., Schaul, T., Silver, D., Hessel, M., Mankowitz, D., Zidek, A., and Munos, R · 2018
Cited alongside, same era.
Reasoning about hypothetical agent behaviours and their parameters
Albrecht, S. V. and Stone, P · 2019
Cited alongside, same era.
Quantifying generalization in reinforcement learning
Cobbe, K., Klimov, O., Hesse, C., Kim, T., and Schulman, J · 2019
Cited alongside, same era.
Wang, T., Gupta, T., Mahajan, A., Peng, B., Whiteson, S., and Zhang, C · 2020
Later among the works it cites.
Multi-task reinforcement learning as a hidden-parameter block mdp
Zhang, A., Sodhani, S., Khetarpal, K., and Pineau, J · 2020
Later among the works it cites.
Why Generalization in RL is Difficult: Epistemic POMDPs and Implicit Partial Observability
Ghosh, D., Rahme, J., Kumar, A., Zhang, A., Adams, R. P., and Levine, S · 2021
Later among the works it cites.
Uneven: Universal value exploration for multi-agent reinforcement learning
Gupta, T., Mahajan, A., Peng, B., Böhmer, W., and Whiteson, S · 2021
Later among the works it cites.
Randomized entity-wise factorization for multi-agent reinforcement learning, 2021
Iqbal, S., de Witt, C. A. S., Peng, B., Böhmer, W., Whiteson, S., and Sha, F · 2021
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels, 2021
Kostrikov, I., Yarats, D., and Fergus, R · 2021
Later among the works it cites.
Scalable evaluation of multi-agent reinforcement learning with melting pot
Leibo, J. Z., Dueñez-Guzman, E. A., Vezhnevets, A., Agapiou, J. P., Sunehag, P., Koster, R., Matyas, J., Beattie, C., Mordatch, I., and Graepel, T · 2021
Later among the works it cites.
Tesseract: Tensorised actors for multi-agent reinforcement learning
Mahajan, A., Samvelyan, M., Mao, L., Makoviychuk, V., Garg, A., Kossaifi, J., Whiteson, S., Zhu, Y., and Anandkumar, A · 2021
Later among the works it cites.
When Is Generalizable Reinforcement Learning Tractable?
Malik, D., Li, Y., and Ravikumar, P · 2021
Later among the works it cites.
Evolutionary dynamics and ϕ \phi -regret minimization in games, 2021
Piliouras, G., Rowland, M., Omidshafiei, S., Elie, R., Hennes, D., Connor, J., and Tuyls, K · 2021
Later among the works it cites.
Automatic data augmentation for generalization in deep reinforcement learning, 2021
Raileanu, R., Goldstein, M., Yarats, D., Kostrikov, I., and Fergus, R · 2021
Later among the works it cites.
Open-ended learning leads to generally capable agents, 2021
Team, O. E. L., Stooke, A., Mahajan, A., Barros, C., Deck, C., Bauer, J., Sygnowski, J., Trebacz, M., Jaderberg, M., Mathieu, M., McAleese, N., Bradley-Schmieg, N., Wong, N., Porcel, N., Raileanu, R., Hughes-Fitt, S., Dalibard, V., and Czarnecki, W. M · 2021
Later among the works it cites.
The surprising effectiveness of ppo in cooperative, multi-agent games, 2021
Yu, C., Velu, A., Vinitsky, E., Wang, Y., Bayen, A., and Wu, Y · 2021
Later among the works it cites.
A survey of generalisation in deep reinforcement learning, 2022
Kirk, R., Zhang, A., Grefenstette, E., and Rocktäschel, T · 2022
Closest in time.