Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine · 2018
Later among the works it cites.
Model-based reinforcement learning via meta-policy optimization
Original
Ignasi Clavera, Jonas Rothfuss, John Schulman, Yasuhiro Fujita, Tamim Asfour, and Pieter Abbeel · 2018
Later among the works it cites.
Probabilistic recurrent state-space models
Original
Andreas Doerr, Christian Daniel, Martin Schiegg, Duy Nguyen-Tuong, Stefan Schaal, Marc Toussaint, and Sebastian Trimpe · 2018
Later among the works it cites.
Visual foresight: Model-based deep reinforcement learning for vision-based robotic control
Original
Frederik Ebert, Chelsea Finn, Sudeep Dasari, Annie Xie, Alex Lee, and Sergey Levine · 2018
Later among the works it cites.
TreeQN and ATreeC: Differentiable tree planning for deep reinforcement learning
Gregory Farquhar, Tim Rocktäschel, Maximilian Igl, and SA Whiteson · 2018
Later among the works it cites.
Model-based value estimation for efficient model-free reinforcement learning
Original
Vladimir Feinberg, Alvin Wan, Ion Stoica, Michael I Jordan, Joseph E Gonzalez, and Sergey Levine · 2018
Later among the works it cites.
Learning to search with MCTSnets
Original
Arthur Guez, Théophane Weber, Ioannis Antonoglou, Karen Simonyan, Oriol Vinyals, Daan Wierstra, Rémi Munos, and David Silver · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Original
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Learning latent dynamics for planning from pixels
Original
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson · 2018
Later among the works it cites.
Model-based reinforcement learning
Jonathan Hui · 2018
Later among the works it cites.
Re-evaluate: Reproducibility in evaluating reinforcement learning algorithms
Khimya Khetarpal, Zafarali Ahmed, Andre Cianflone, Riashat Islam, and Joelle Pineau · 2018
Later among the works it cites.
Ranked reward: Enabling self-play reinforcement learning for combinatorial optimization
Original
Alexandre Laterre, Yunguan Fu, Mohamed Khalil Jabri, Alain-Sam Cohen, David Kas, Karl Hajjar, Torbjorn S Dahl, Amine Kerkeni, and Karim Beguir · 2018
Later among the works it cites.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Marlos C Machado, Marc G Bellemare, Erik Talvitie, Joel Veness, Matthew Hausknecht, and Michael Bowling · 2018
Later among the works it cites.
A0c: Alpha zero in continuous action space
Original
Thomas M Moerland, Joost Broekens, Aske Plaat, and Catholijn M Jonker · 2018
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Anusha Nagabandi, Gregory Kahn, Ronald S Fearing, and Sergey Levine · 2018
Later among the works it cites.
Value propagation networks
Original
Nantas Nardelli, Gabriel Synnaeve, Zeming Lin, Pushmeet Kohli, Philip HS Torr, and Nicolas Usunier · 2018
Later among the works it cites.
Generalized value iteration networks: Life beyond lattices
Sufeng Niu, Siheng Chen, Hanyu Guo, Colin Targonski, Melissa C Smith, and Jelena Kovačević · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis · 2018
Later among the works it cites.
Universal planning networks
Original
Aravind Srinivas, Allan Jabri, Pieter Abbeel, Sergey Levine, and Chelsea Finn · 2018
Later among the works it cites.
Reinforcement learning, An Introduction, Second Edition
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Deepmind control suite
Original
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al · 2018
Later among the works it cites.
Deep reinforcement learning for general video game ai
Ruben Rodriguez Torrado, Philip Bontrager, Julian Togelius, Jialin Liu, and Diego Perez-Liebana · 2018
Later among the works it cites.
Evolving mario levels in the latent space of a deep convolutional generative adversarial network
Vanessa Volz, Jacob Schrum, Jialin Liu, Simon M Lucas, Adam Smith, and Sebastian Risi · 2018
Later among the works it cites.
Meta-learning: Learning to learn fast
Lilian Weng · 2018
Later among the works it cites.
Relational deep reinforcement learning
Original
Vinicius Zambaldi, David Raposo, Adam Santoro, Victor Bapst, Yujia Li, Igor Babuschkin, Karl Tuyls, David Reichert, Timothy Lillicrap, Edward Lockhart, et al · 2018
Later among the works it cites.
Policy gradient search: Online planning and expert iteration without search trees
Original
Thomas Anthony, Robert Nishihara, Philipp Moritz, Tim Salimans, and John Schulman · 2019
Later among the works it cites.
Dota 2 with large scale deep reinforcement learning
Original
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemysław Debiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, et al · 2019
Later among the works it cites.
Superhuman ai for multiplayer poker
Noam Brown and Tuomas Sandholm · 2019
Later among the works it cites.
Model-free reinforcement learning algorithms: A survey
Sinan Çalışır and Meltem Kurt Pehlivanoğlu · 2019
Later among the works it cites.
On-line adaptative curriculum learning for gans
Thang Doan, Joao Monteiro, Isabela Albuquerque, Bogdan Mazoure, Audrey Durand, Joelle Pineau, and R Devon Hjelm · 2019
Later among the works it cites.
An investigation of model-free planning
Original
Arthur Guez, Mehdi Mirza, Karol Gregor, Rishabh Kabra, Sébastien Racanière, Théophane Weber, David Raposo, Adam Santoro, Laurent Orseau, Tom Eccles, et al · 2019
Later among the works it cites.
Dream to control: Learning behaviors by latent imagination
Original
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi · 2019
Later among the works it cites.
Analogues of mental simulation and imagination in deep learning
Jessica B Hamrick · 2019
Later among the works it cites.
Hex: The Full Story
Ryan B Hayward and Bjarne Toft · 2019
Later among the works it cites.
Human-level performance in 3D multiplayer games with population-based reinforcement learning
Max Jaderberg, Wojciech M Czarnecki, Iain Dunning, Luke Marris, Guy Lever, Antonio Garcia Castaneda, Charles Beattie, Neil C Rabinowitz, Ari S Morcos, Avraham Ruderman, et al · 2019
Later among the works it cites.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Later among the works it cites.
Obstacle tower: A generalization challenge in vision, control, and planning
Original
Arthur Juliani, Ahmed Khalifa, Vincent-Pierre Berges, Jonathan Harper, Ervin Teng, Hunter Henry, Adam Crespi, Julian Togelius, and Danny Lange · 2019
Later among the works it cites.
Deep learning for video game playing
Niels Justesen, Philip Bontrager, Julian Togelius, and Sebastian Risi · 2019
Later among the works it cites.
Model-based reinforcement learning for Atari
Original
Lukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski, Roy H Campbell, Konrad Czechowski, Dumitru Erhan, Chelsea Finn, Piotr Kozakowski, Sergey Levine, et al · 2019
Later among the works it cites.
An introduction to variational autoencoders
Original
Diederik P Kingma and Max Welling · 2019
Later among the works it cites.
Emergent coordination through competition
Original
Siqi Liu, Guy Lever, Josh Merel, Saran Tunyasuvunakool, Nicolas Heess, and Thore Graepel · 2019
Later among the works it cites.
Deep neuroevolution of recurrent and discrete world models
Sebastian Risi and Kenneth O Stanley · 2019
Later among the works it cites.
Value iteration networks on multiple levels of abstraction
Original
Daniel Schleich, Tobias Klamt, and Sven Behnke · 2019
Later among the works it cites.
Mastering Atari, Go, chess and shogi by planning with a learned model
Original
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Later among the works it cites.
Alternative loss functions in AlphaZero-like self-play
Hui Wang, Michael Emmerich, Mike Preuss, and Aske Plaat · 2019
Later among the works it cites.
Introduction to machine learning, Third edition
Ethem Alpaydin · 2020
Closest in time.
Learning to play no-press diplomacy with best response policy iteration
Original
Thomas Anthony, Tom Eccles, Andrea Tacchetti, János Kramár, Ian Gemp, Thomas C Hudson, Nicolas Porcel, Marc Lanctot, Julien Pérolat, Richard Everett, et al · 2020
Closest in time.
Agent57: Outperforming the Atari human benchmark
Original
Adrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann, Alex Vitvitskyi, Daniel Guo, and Charles Blundell · 2020
Closest in time.
Solving hard AI planning instances using curriculum-driven deep reinforcement learning
Original
Dieqiao Feng, Carla P Gomes, and Bart Selman · 2020
Closest in time.
Meta-learning in neural networks: A survey
Original
Timothy Hospedales, Antreas Antoniou, Paul Micaelli, and Amos Storkey · 2020
Closest in time.
Metalearning for deep neural networks
Mike Huisman, Jan van Rijn, and Aske Plaat · 2020
Closest in time.
Curriculum learning for reinforcement learning domains: A framework and survey
Original
Sanmit Narvekar, Bei Peng, Matteo Leonetti, Jivko Sinapov, Matthew E Taylor, and Peter Stone · 2020
Closest in time.
Learning to Play: Reinforcement Learning and Games
Aske Plaat · 2020
Closest in time.
From Chess and Atari to StarCraft and Beyond: How Game AI is Driving the World of AI
Sebastian Risi and Mike Preuss · 2020
Closest in time.
Planning to explore via self-supervised world models
Original
Ramanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel, Danijar Hafner, and Deepak Pathak · 2020
Closest in time.
Tackling Morpion Solitaire with AlphaZero-like Ranked Reward reinforcement learning
Original
Hui Wang, Mike Preuss, Michael Emmerich, and Aske Plaat · 2020
Closest in time.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Karol Hausman, Chelsea Finn, and Sergey Levine · 2020
Closest in time.