Fetching the paper…
Reading the bibliography…
Model-free reinforcement learning (RL) can be used to learn effective policies for complex tasks, such as Atari games, even from image observations.
Dyna, an integrated architecture for learning, planning, and reacting
Richard S. Sutton · 1991
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990-2010)
Jürgen Schmidhuber · 2010
Earlier work this paper cites.
A survey on policy search for robotics
Marc Peter Deisenroth, Gerhard Neumann, and Jan Peters · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller · 2013
Earlier work this paper cites.
Learning behaviors of and interactions among objects through spatio–temporal reasoning
Mustafa Ersen and Sanem Sariel · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents (extended abstract)
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2015
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer · 2015
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Timothy P. Lillicrap, Tom Erez, and Yuval Tassa · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin A. Riedmiller, Andreas Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard L. Lewis, and Satinder P. Singh · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Manuel Watter, Jost Tobias Springenberg, Joschka Boedecker, and Martin A. Riedmiller · 2015
Earlier work this paper cites.
Deep spatial autoencoders for visuomotor learning
Chelsea Finn, Xin Yu Tan, Yan Duan, Trevor Darrell, Sergey Levine, and Pieter Abbeel · 2016
Earlier work this paper cites.
A deep learning approach for joint video frame and reward prediction in Atari games
Felix Leibfried, Nate Kushman, and Katja Hofmann · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Improved learning of dynamics models for control
Arun Venkatraman, Roberto Capobianco, Lerrel Pinto, Martial Hebert, Daniele Nardi, and J. Andrew Bagnell · 2016
Earlier work this paper cites.
Reinforcement learning through asynchronous advantage actor-critic on a GPU
Mohammad Babaeizadeh, Iuri Frosio, Stephen Tyree, Jason Clemons, and Jan Kautz · 2017
Cited alongside, same era.
Recurrent environment simulators
Silvia Chiappa, Sébastien Racanière, Daan Wierstra, and Shakir Mohamed · 2017
Cited alongside, same era.
Self-supervised visual planning with temporal skip connections
Frederik Ebert, Chelsea Finn, Alex X. Lee, and Sergey Levine · 2017
Cited alongside, same era.
Deep visual foresight for planning robot motion
Chelsea Finn and Sergey Levine · 2017
Cited alongside, same era.
Game engine learning from video
Matthew Guzdial, Boyang Li, and Mark O. Riedl · 2017
Cited alongside, same era.
Uncertainty-driven imagination for continuous deep reinforcement learning
Gabriel Kalweit and Joschka Boedecker · 2017
Cited alongside, same era.
Model-based value estimation for efficient model-free reinforcement learning
Vladimir Feinberg, Alvin Wan, Ion Stoica, Michael I. Jordan, Joseph E. Gonzalez, and Sergey Levine · 2018
Later among the works it cites.
Recurrent world models facilitate policy evolution
David Ha and Jürgen Schmidhuber · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Gheshlaghi Azar, and David Silver · 2018
Later among the works it cites.
The effect of planning shape on dyna-style planning in high-dimensional state spaces
G. Zacharias Holland, Erik Talvitie, and Michael Bowling · 2018
Later among the works it cites.
Discrete autoencoders for sequence models
Lukasz Kaiser and Samy Bengio · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Value prediction network
Junhyuk Oh, Satinder Singh, and Honglak Lee · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Reinforcement learning - an introduction, 2nd edition (work in progress)
Richard S. Sutton and Andrew G. Barto · 2017
Cited alongside, same era.
Human learning in atari
Pedro Tsividis, Thomas Pouncy, Jaqueline L. Xu, Joshua B. Tenenbaum, and Samuel J. Gershman · 2017
Cited alongside, same era.
Neural discrete representation learning
Aäron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu · 2017
Cited alongside, same era.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Yuhuai Wu, Elman Mansimov, Roger B. Grosse, Shun Liao, and Jimmy Ba · 2017
Cited alongside, same era.
Later among the works it cites.
Model-ensemble trust-region policy optimization
Thanard Kurutach, Ignasi Clavera, Yan Duan, Aviv Tamar, and Pieter Abbeel · 2018
Later among the works it cites.
Model-based regularization for deep reinforcement learning with transcoder networks
Felix Leibfried, Rasul Tutunov, Peter Vrancx, and Haitham Bou-Ammar · 2018
Later among the works it cites.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Marlos C. Machado, Marc G. Bellemare, Erik Talvitie, Joel Veness, Matthew J. Hausknecht, and Michael Bowling · 2018
Later among the works it cites.
Learning real-world robot policies by dreaming
A. J. Piergiovanni, Alan Wu, and Michael S. Ryoo · 2018
Later among the works it cites.
Observe and look further: Achieving consistent performance on atari
Tobias Pohlen, Bilal Piot, Todd Hester, Mohammad Gheshlaghi Azar, Dan Horgan, David Budden, Gabriel Barth-Maron, Hado van Hasselt, John Quan, Mel Vecerík, Matteo Hessel, Rémi Munos, and Olivier Pietquin · 2018
Later among the works it cites.
Unsupervised learning of sensorimotor affordances by stochastic future prediction
Oleh Rybkin, Karl Pertsch, Andrew Jaegle, Konstantinos G. Derpanis, and Kostas Daniilidis · 2018
Later among the works it cites.
Woulda, coulda, shoulda: Counterfactually-guided policy search
Lars Buesing, Theophane Weber, Yori Zwols, Nicolas Heess, Sébastien Racanière, Arthur Guez, and Jean-Baptiste Lespiau · 2019
Closest in time.
Learning latent dynamics for planning from pixels
Danijar Hafner, Timothy P. Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson · 2019
Closest in time.
Visual robot task planning
Chris Paxton, Yotam Barnoy, Kapil D. Katyal, Raman Arora, and Gregory D. Hager · 2019
Closest in time.
Learning powerful policies by using consistent dynamics model
Shagun Sodhani, Anirudh Goyal, Tristan Deleu, Yoshua Bengio, Sergey Levine, and Jian Tang · 2019
Closest in time.
When to use parametric models in reinforcement learning?
Hado van Hasselt, Matteo Hessel, and John Aslanides · 2019
Closest in time.
Benchmarking model-based reinforcement learning
Tingwu Wang, Xuchan Bao, Ignasi Clavera, Jerrick Hoang, Yeming Wen, Eric Langlois, Shunshi Zhang, Guodong Zhang, Pieter Abbeel, and Jimmy Ba · 2019
Closest in time.
Do recent advancements in model-based deep reinforcement learning really improve data efficiency?, 2020
Kacper Piotr Kielak · 2020
Closest in time.