Open-ended learning in symmetric zero-sum games
D. Balduzzi, M. Garnelo, Y. Bachrach, W. Czarnecki, J. Perolat, M. Jaderberg, and T. Graepel · 2019
Later among the works it cites.
Dota 2 with large scale deep reinforcement learning
Original
C. Berner, G. Brockman, B. Chan, V. Cheung, P. Dębiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, et al · 2019
Later among the works it cites.
AI-GAs: AI-generating algorithms, an alternate paradigm for producing general artificial intelligence
Original
J. Clune · 2019
Later among the works it cites.
CURIOUS: intrinsically motivated modular multi-goal reinforcement learning
C. Colas, P. Oudeyer, O. Sigaud, P. Fournier, and M. Chetouani · 2019
Later among the works it cites.
Distilling policy distillation
W. M. Czarnecki, R. Pascanu, S. Osindero, S. Jayakumar, G. Swirszcz, and M. Jaderberg · 2019
Later among the works it cites.
Curriculum-guided hindsight experience replay
M. Fang, T. Zhou, Y. Du, L. Han, and Z. Zhang · 2019
Later among the works it cites.
Multi-task deep reinforcement learning with popart
M. Hessel, H. Soyer, L. Espeholt, W. Czarnecki, S. Schmitt, and H. van Hasselt · 2019
Later among the works it cites.
Unsupervised curricula for visual meta-reinforcement learning
A. Jabri, K. Hsu, A. Gupta, B. Eysenbach, S. Levine, and C. Finn · 2019
Later among the works it cites.
Human-level performance in 3d multiplayer games with population-based reinforcement learning
M. Jaderberg, W. M. Czarnecki, I. Dunning, L. Marris, G. Lever, A. G. Castaneda, C. Beattie, N. C. Rabinowitz, A. S. Morcos, A. Ruderman, et al · 2019
Later among the works it cites.
MAVEN: Multi-agent variational exploration
A. Mahajan, T. Rashid, M. Samvelyan, and S. Whiteson · 2019
Later among the works it cites.
Teacher–student curriculum learning
T. Matiisen, A. Oliver, T. Cohen, and J. Schulman · 2019
Later among the works it cites.
Composing value functions in reinforcement learning
B. Van Niekerk, S. James, A. Earle, and B. Rosman · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al · 2019
Later among the works it cites.
POET: open-ended coevolution of environments and their optimized solutions
R. Wang, J. Lehman, J. Clune, and K. O. Stanley · 2019
Later among the works it cites.
Unsupervised control through non-parametric discriminative rewards
D. Warde-Farley, T. V. de Wiele, T. D. Kulkarni, C. Ionescu, S. Hansen, and V. Mnih · 2019
Later among the works it cites.
Emergent tool use from multi-agent autocurricula
B. Baker, I. Kanitscheider, T. M. Markov, Y. Wu, G. Powell, B. McGrew, and I. Mordatch · 2020
Later among the works it cites.
Language-conditioned goal generation: a new approach to language grounding for RL
Original
C. Colas, A. Akakzia, P.-Y. Oudeyer, M. Chetouani, and O. Sigaud · 2020
Later among the works it cites.
Real world games look like spinning tops
W. M. Czarnecki, G. Gidel, B. Tracey, K. Tuyls, S. Omidshafiei, D. Balduzzi, and M. Jaderberg · 2020
Later among the works it cites.
Emergent complexity and zero-shot transfer via unsupervised environment design
M. Dennis, N. Jaques, E. Vinitsky, A. Bayen, S. Russell, A. Critch, and S. Levine · 2020
Later among the works it cites.
A boolean task algebra for reinforcement learning
G. Nangue Tasse, S. James, and B. Rosman · 2020
Later among the works it cites.
Skew-fit: State-covering self-supervised reinforcement learning
V. Pong, M. Dalal, S. Lin, A. Nair, S. Bahl, and S. Levine · 2020
Later among the works it cites.
Automatic curriculum learning for deep RL: A short survey
R. Portelas, C. Colas, L. Weng, K. Hofmann, and P.-Y. Oudeyer · 2020
Later among the works it cites.
Automated curriculum generation through setter-solver interactions
S. Racanière, A. K. Lampinen, A. Santoro, D. P. Reichert, V. Firoiu, and T. P. Lillicrap · 2020
Later among the works it cites.
CPPN2GAN: Combining compositional pattern producing networks and gans for large-scale pattern generation
J. Schrum, V. Volz, and S. Risi · 2020
Later among the works it cites.
Contextual games: Multi-agent learning with side information
P. G. Sessa, I. Bogunovic, A. Krause, and M. Kamgarpour · 2020
Later among the works it cites.
V-MPO: on-policy maximum a posteriori policy optimization for discrete and continuous control
H. F. Song, A. Abdolmaleki, J. T. Springenberg, A. Clark, H. Soyer, J. W. Rae, S. Noury, A. Ahuja, S. Liu, D. Tirumala, N. Heess, D. Belov, M. A. Riedmiller, and M. M. Botvinick · 2020
Later among the works it cites.
Options as responses: Grounding behavioural hierarchies in multi-agent reinforcement learning
A. Vezhnevets, Y. Wu, M. Eckstein, R. Leblond, and J. Z. Leibo · 2020
Later among the works it cites.
Enhanced POET: open-ended reinforcement learning through unbounded invention of learning challenges and their solutions
R. Wang, J. Lehman, A. Rawal, J. Zhi, Y. Li, J. Clune, and K. O. Stanley · 2020
Later among the works it cites.
Automatic curriculum learning through value disagreement
Y. Zhang, P. Abbeel, and L. Pinto · 2020
Later among the works it cites.
Learning with amigo: Adversarially motivated intrinsic goals
A. Campero, R. Raileanu, H. Küttler, J. B. Tenenbaum, T. Rocktäschel, and E. Grefenstette · 2021
Closest in time.
Pick your battles: Interaction graphs as population-level objectives for strategic diversity
M. Garnelo, W. M. Czarnecki, S. Liu, D. Tirumala, J. Oh, G. Gidel, H. van Hasselt, and D. Balduzzi · 2021
Closest in time.
Multimodal neurons in artificial neural networks
G. Goh, N. C. †, C. V. †, S. Carter, M. Petrov, L. Schubert, A. Radford, and C. Olah · 2021
Closest in time.
Multi-task curriculum learning in a complex, visual, hard-exploration domain: Minecraft
Original
I. Kanitscheider, J. Huizinga, D. Farhi, W. H. Guss, B. Houghton, R. Sampedro, P. Zhokhov, B. Baker, A. Ecoffet, J. Tang, O. Klimov, and J. Clune · 2021
Closest in time.
Multi-agent training beyond zero-sum with correlated equilibrium meta-solvers
L. Marris, P. Muller, M. Lanctot, K. Tuyls, and T. Graepel · 2021
Closest in time.
A graph placement methodology for fast chip design
A. Mirhoseini, A. Goldie, M. Yazgan, J. W. Jiang, E. Songhori, S. Wang, Y.-J. Lee, E. Johnson, O. Pathak, A. Nazi, et al · 2021
Closest in time.
Asymmetric self-play for automatic goal discovery in robotic manipulation
Original
OpenAI, M. Plappert, R. Sampedro, T. Xu, I. Akkaya, V. Kosaraju, P. Welinder, R. D’Sa, A. Petron, H. P. de Oliveira Pinto, A. Paino, H. Noh, L. Weng, Q. Yuan, C. Chu, and W. Zaremba · 2021
Closest in time.