Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Later among the works it cites.
The option-critic architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup · 2017
Later among the works it cites.
Recurrent environment simulators
Original
Silvia Chiappa, Sébastien Racaniere, Daan Wierstra, and Shakir Mohamed · 2017
Later among the works it cites.
Z-forcing: Training stochastic recurrent networks
Anirudh ALIAS PARTH Goyal, Alessandro Sordoni, Marc-Alexandre Côté, Nan Ke, and Yoshua Bengio · 2017
Later among the works it cites.
Unsupervised real-time control through variational empowerment
Original
Maximilian Karl, Maximilian Soelch, Philip Becker-Ehmck, Djalel Benbouzid, Patrick van der Smagt, and Justin Bayer · 2017
Later among the works it cites.
Prediction and control with temporal segment models
Original
Nikhil Mishra, Pieter Abbeel, and Igor Mordatch · 2017
Later among the works it cites.
Value prediction network
Junhyuk Oh, Satinder Singh, and Honglak Lee · 2017
Later among the works it cites.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Original
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Feudal networks for hierarchical reinforcement learning
Original
Alexander Sasha Vezhnevets, Simon Osindero, Tom Schaul, Nicolas Heess, Max Jaderberg, David Silver, and Koray Kavukcuoglu · 2017
Later among the works it cites.
Lipschitz continuity in model-based reinforcement learning
Original
Kavosh Asadi, Dipendra Misra, and Michael L Littman · 2018
Later among the works it cites.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Original
Jacob Buckman, Danijar Hafner, George Tucker, Eugene Brevdo, and Honglak Lee · 2018
Later among the works it cites.
Learning and querying fast generative models for reinforcement learning
Original
Lars Buesing, Theophane Weber, Sebastien Racaniere, SM Eslami, Danilo Rezende, David P Reichert, Fabio Viola, Frederic Besse, Karol Gregor, Demis Hassabis, et al · 2018
Later among the works it cites.
Minimalistic gridworld environment for openai gym
Maxime Chevalier-Boisvert and Lucas Willems · 2018
Later among the works it cites.
Self-consistent trajectory autoencoder: Hierarchical reinforcement learning with trajectory embeddings
Original
John D Co-Reyes, YuXuan Liu, Abhishek Gupta, Benjamin Eysenbach, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Generating sentences by editing prototypes
Kelvin Guu, Tatsunori B Hashimoto, Yonatan Oren, and Percy Liang · 2018
Later among the works it cites.
Recurrent world models facilitate policy evolution
Original
David Ha and Jürgen Schmidhuber · 2018
Later among the works it cites.
The effect of planning shape on dyna-style planning in high-dimensional state spaces
Original
G Zacharias Holland, Erik Talvitie, and Michael Bowling · 2018
Later among the works it cites.