Openai five
OpenAI · 2018
Later among the works it cites.
Taming vaes
Original
Danilo Jimenez Rezende and Fabio Viola · 2018
Later among the works it cites.
Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
Original
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis · 2018
Later among the works it cites.
Solving the rubik’s cube with deep reinforcement learning and search
Forest Agostinelli, Stephen McAleer, Alexander Shmakov, and Pierre Baldi · 2019
Closest in time.
Human-level performance in 3d multiplayer games with population-based reinforcement learning
Max Jaderberg, Wojciech M. Czarnecki, Iain Dunning, Luke Marris, Guy Lever, Antonio Garcia Castañeda, Charles Beattie, Neil C. Rabinowitz, Ari S. Morcos, Avraham Ruderman, Nicolas Sonnerat, Tim Green, Louise Deason, Joel Z. Leibo, David Silver, Demis Hassabis, Koray Kavukcuoglu, and Thore Graepel · 2019
Closest in time.
Biva: A very deep hierarchy of latent variables for generative modeling
Original
Lars Maaløe, Marco Fraccaro, Valentin Liévin, and Ole Winther · 2019
Closest in time.
Learning curriculum policies for reinforcement learning
Sanmit Narvekar and Peter Stone · 2019
Closest in time.
Classification accuracy score for conditional generative models
Original
Suman Ravuri and Oriol Vinyals · 2019
Closest in time.
AlphaStar: Mastering the Real-Time Strategy Game StarCraft II
Oriol Vinyals, Igor Babuschkin, Junyoung Chung, Michael Mathieu, Max Jaderberg, Wojciech M. Czarnecki, Andrew Dudzik, Aja Huang, Petko Georgiev, Richard Powell, Timo Ewalds, Dan Horgan, Manuel Kroiss, Ivo Danihelka, John Agapiou, Junhyuk Oh, Valentin Dalibard, David Choi, Laurent Sifre, Yury Sulsky, Sasha Vezhnevets, James Molloy, Trevor Cai, David Budden, Tom Paine, Caglar Gulcehre, Ziyu Wang, Tobias Pfaff, Toby Pohlen, Yuhuai Wu, Dani Yogatama, Julia Cohen, Katrina McKinney, Oliver Smith, Tom Schaul, Timothy Lillicrap, Chris Apps, Koray Kavukcuoglu, Demis Hassabis, and David Silver · 2019
Closest in time.
Paired open-ended trailblazer (poet): Endlessly generating increasingly complex and diverse learning environments and their solutions
Original
Rui Wang, Joel Lehman, Jeff Clune, and Kenneth O Stanley · 2019
Closest in time.
Deep reinforcement learning with relational inductive biases
Vinicius Zambaldi, David Raposo, Adam Santoro, Victor Bapst, Yujia Li, Igor Babuschkin, Karl Tuyls, David Reichert, Timothy Lillicrap, Edward Lockhart, Murray Shanahan, Victoria Langston, Razvan Pascanu, Matthew Botvinick, Oriol Vinyals, and Peter Battaglia · 2019
Closest in time.