Fetching the paper…
Reading the bibliography…
We introduce Imagination-Augmented Agents (I2As), a novel architecture for deep reinforcement learning combining model-free and model-based aspects.
Cognitive maps in rats and men
Edward C Tolman · 1948
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton · 1990
Earlier work this paper cites.
An on-line algorithm for dynamic reinforcement learning and planning in reactive environments
Jürgen Schmidhuber · 1990
Earlier work this paper cites.
Efficient learning and planning within the dyna framework
Jing Peng and Ronald J Williams · 1993
Earlier work this paper cites.
Advantage updating
Leemon C Baird III · 1993
Earlier work this paper cites.
On-line policy improvement using monte-carlo search
Gerald Tesauro and Gregory R Galperin · 1996
Earlier work this paper cites.
Automatic making of sokoban problems
Yoshio Murase, Hitoshi Matsubara, and Yuzuru Hiraga · 1996
Earlier work this paper cites.
The Role of Learning in the Operation of Motivational Systems
Anthony Dickinson and Bernard Balleine · 2002
Earlier work this paper cites.
Exploration and apprenticeship learning in reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2005
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
Rémi Coulom · 2006
Earlier work this paper cites.
Bandit based monte-carlo planning
Levente Kocsis and Csaba Szepesvári · 2006
Earlier work this paper cites.
Universal intelligence: A definition of machine intelligence
Shane Legg and Marcus Hutter · 2007
Earlier work this paper cites.
Combining online and offline knowledge in uct
Sylvain Gelly and David Silver · 2007
Earlier work this paper cites.
Transpositions and move groups in monte carlo tree search
Benjamin E Childs, James H Brodeur, and Levente Kocsis · 2008
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
Matthew E Taylor and Peter Stone · 2009
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Marc Deisenroth and Carl E Rasmussen · 2011
Earlier work this paper cites.
Nested rollout policy adaptation for monte carlo tree search
Christopher D Rosin · 2011
Earlier work this paper cites.
Procedural generation of sokoban levels
Joshua Taylor and Ian Parberry · 2011
Earlier work this paper cites.
The future of memory: remembering, imagining, and the brain
Daniel L Schacter, Donna Rose Addis, Demis Hassabis, Victoria C Martin, R Nathan Spreng, and Karl K Szpunar · 2012
Cited alongside, same era.
Lecture 6.5-RMSprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Hippocampal place-cell sequences depict future paths to remembered goals
Brad E Pfeiffer and David J Foster · 2013
Cited alongside, same era.
Learning neural network policies with guided policy search under unknown dynamics
Sergey Levine and Pieter Abbeel · 2014
Cited alongside, same era.
Transfer from simulation to real world through learning deep inverse dynamics model
Paul Christiano, Zain Shah, Igor Mordatch, Jonas Schneider, Trevor Blackwell, Joshua Tobin, Pieter Abbeel, and Wojciech Zaremba · 2016
Later among the works it cites.
Continuous deep q-learning with model-based acceleration
Shixiang Gu, Timothy Lillicrap, Ilya Sutskever, and Sergey Levine · 2016
Later among the works it cites.
Improved learning of dynamics models for control
Arun Venkatraman, Roberto Capobianco, Lerrel Pinto, Martial Hebert, Daniele Nardi, and J Andrew Bagnell · 2016
Later among the works it cites.
Value iteration networks
Aviv Tamar, Yi Wu, Garrett Thomas, Sergey Levine, and Pieter Abbeel · 2016
Later among the works it cites.
Reinforcement learning with unsupervised auxiliary tasks
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z Leibo, David Silver, and Koray Kavukcuoglu · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Model regularization for stable sample rollouts
Erik Talvitie · 2014
Cited alongside, same era.
Agnostic system identification for monte carlo planning
Erik Talvitie · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard L Lewis, and Satinder Singh · 2015
Cited alongside, same era.
Embed to control: A locally linear latent dynamics model for control from raw images
Manuel Watter, Jost Springenberg, Joschka Boedecker, and Martin Riedmiller · 2015
Cited alongside, same era.
DeepMPC: Learning deep latent features for model predictive control
Ian Lenz, Ross A Knepper, and Ashutosh Saxena · 2015
Cited alongside, same era.
Towards adapting deep visuomotor representations from simulated to real environments
Eric Tzeng, Coline Devin, Judy Hoffman, Chelsea Finn, Xingchao Peng, Sergey Levine, Kate Saenko, and Trevor Darrell · 2015
Cited alongside, same era.
Scheduled sampling for sequence prediction with recurrent neural networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer · 2015
Cited alongside, same era.
Learning to navigate in complex environments
Piotr Mirowski, Razvan Pascanu, Fabio Viola, Hubert Soyer, Andy Ballard, Andrea Banino, Misha Denil, Ross Goroshin, Laurent Sifre, Koray Kavukcuoglu, et al · 2016
Later among the works it cites.
Adaptive computation time for recurrent neural networks
Alex Graves · 2016
Later among the works it cites.
Recurrent environment simulators
Silvia Chiappa, Sébastien Racaniere, Daan Wierstra, and Shakir Mohamed · 2017
Closest in time.
https://drive.google.com/open?id=0B4tKsKnCCZtQY2tTOThucHVxUTQ , 2017
2017
Closest in time.
Deep visual foresight for planning robot motion
Chelsea Finn and Sergey Levine · 2017
Closest in time.
Imitation from observation: Learning to imitate behaviors from raw video via context translation
YuXuan Liu, Abhishek Gupta, Pieter Abbeel, and Sergey Levine · 2017
Closest in time.
Goal-driven dynamics learning via bayesian optimization
Somil Bansal, Roberto Calandra, Ted Xiao, Sergey Levine, and Claire J Tomlin · 2017
Closest in time.
Alonso Marco, Felix Berkenkamp, Philipp Hennig, Angela P Schoellig, Andreas Krause, Stefan Schaal, and Sebastian Trimpe · 2017
Closest in time.
Junhyuk Oh, Satinder Singh, and Honglak Lee · 2017
Closest in time.
Model-based planning in discrete action spaces
Mikael Henaff, William F Whitney, and Yann LeCun · 2017
Closest in time.
Schema networks: Zero-shot transfer with a generative causal model of intuitive physics
Ken Kansky, Tom Silver, David A Mély, Mohamed Eldawy, Miguel Lázaro-Gredilla, Xinghua Lou, Nimrod Dorfman, Szymon Sidor, Scott Phoenix, and Dileep George · 2017
Closest in time.
Metacontrol for adaptive imagination-based optimization
Jessica B. Hamrick, Andy J. Ballard, Razvan Pascanu, Oriol Vinyals, Nicolas Heess, and Peter W. Battaglia · 2017
Closest in time.
Learning model-based planning from scratch
Razvan Pascanu, Yujia Li, Oriol Vinyals, Nicolas Heess, David Reichert, Theophane Weber, Sebastien Racaniere, Lars Buesing, Daan Wierstra, and Peter Battaglia · 2017
Closest in time.