Fetching the paper…
Reading the bibliography…
Conventional wisdom holds that model-based planning is a powerful approach to sequential decision-making.
Reasoning about beliefs and actions under computational resource constraints
Eric J. Horvitz · 1988
Earlier work this paper cites.
Principles of metareasoning
Stuart Russell and Eric Wefald · 1991
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Richard S Sutton · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
Rémi Coulom · 2006
Earlier work this paper cites.
Optimistic planning of deterministic systems
Jean-Francois Hren and Rémi Munos · 2008
Earlier work this paper cites.
Reinforcement learning and dynamic programming using function approximators
Lucian Busoniu, Robert Babuska, Bart De Schutter, and Damien Ernst · 2010
Earlier work this paper cites.
Optimistic optimization of a deterministic function without the knowledge of its smoothness
Rémi Munos · 2011
Earlier work this paper cites.
Selecting computations: Theory and applications
Nicholas Hay, Stuart J. Russell, David Tolpin, and Solomon Eyal Shimony · 2012
Cited alongside, same era.
Bandit-based planning and learning in continuous-action markov decision processes
Ari Weinstein and Michael L Littman · 2012
Cited alongside, same era.
Deep learning of representations: Looking forward
Yoshua Bengio · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Conditional computation in neural networks for faster models
Emmanuel Bengio, Pierre-Luc Bacon, Joelle Pineau, and Doina Precup · 2015
Cited alongside, same era.
Interaction networks for learning about objects, relations and physics
Peter Battaglia, Razvan Pascanu, Matthew Lai, Danilo Jimenez Rezende, et al · 2016
Later among the works it cites.
Adaptive computation time for recurrent neural networks
Alex Graves · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Later among the works it cites.
The predictron: End-to-end learning and planning
David Silver, Hado van Hasselt, Matteo Hessel, Tom Schaul, Arthur Guez, Tim Harley, Gabriel Dulac-Arnold, David Reichert, Neil Rabinowitz, Andre Barreto, et al · 2016
Later among the works it cites.
Value Iteration Networks
Aviv Tamar, Sergey Levine, and Pieter Abbeel · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning Visual Predictive Models of Physics for Playing Billiards
Katerina Fragkiadaki, Pulkit Agrawal, Sergey Levine, and Jitendra Malik · 2015
Cited alongside, same era.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Tim Lillicrap, Tom Erez, and Yuval Tassa · 2015
Cited alongside, same era.
DeepMPC: Learning deep latent features for model predictive control
Ian Lenz, Ross A Knepper, and Ashutosh Saxena · 2015
Cited alongside, same era.
Embed to control: A locally linear latent dynamics model for control from raw images
Manuel Watter, Jost Springenberg, Joschka Boedecker, and Martin Riedmiller · 2015
Cited alongside, same era.
Strategic attentive writer for learning macro-actions
Alexander Vezhnevets, Volodymyr Mnih, John Agapiou, Simon Osindero, Alex Graves, Oriol Vinyals, Koray Kavukcuoglu, et al · 2016
Later among the works it cites.
Deep visual foresight for planning robot motion
Chelsea Finn and Sergey Levine · 2017
Closest in time.
Metacontrol for adaptive imagination-based optimization, 2017
Jessica B. Hamrick, Andrew J. Ballard, Razvan Pascanu, Oriol Vinyals, Nicolas Heess, and Peter W. Battaglia · 2017
Closest in time.
Imagination-augmented agents for deep reinforcement learning
Theophane Weber, Sebastien Racaniere, David P. Reichert, Lars Buesing, Arthur Guez, Danilo Rezende, Adria Puigdomènech Badia, Oriol Vinyals, Nicolas Heess, Yujia Li, Razvan Pascanu, Peter Battaglia, David Silver, and Daan Wierstra · 2017
Closest in time.