Fetching the paper…
Reading the bibliography…
Planning methods can solve temporally extended sequential decision making problems by composing simple behaviors.
Neural networks for self-learning control systems
Derrick H Nguyen and Bernard Widrow · 1990
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton · 1990
Earlier work this paper cites.
Feudal reinforcement learning
Peter Dayan and Geoffrey E Hinton · 1993
Earlier work this paper cites.
Hq-learning
Marco Wiering and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Image manifolds
Haw-Minn Lu, Yeshaiahu Fainman, and Robert Hecht-Nielsen · 1998
Earlier work this paper cites.
Multi-value-functions: effcient automatic action hierarchies for multiple goal mdps
Andrew W Moore, Leemon Baird, and Leslie P Kaelbling · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Thomas G Dietterich · 2000
Earlier work this paper cites.
Toward hierarchical decomposition for planning in uncertain environments
Terran Lane and Leslie Pack Kaelbling · 2001
Earlier work this paper cites.
Structure in the space of value functions
David Foster and Peter Dayan · 2002
Earlier work this paper cites.
A tutorial on the cross-entropy method
Pieter-Tjerk De Boer, Dirk P Kroese, Shie Mannor, and Reuven Y Rubinstein · 2005
Earlier work this paper cites.
Numerical optimization
Jorge Nocedal and Stephen Wright · 2006
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Richard S Sutton, Joseph Modayil, Michael Delp, Thomas Degris, Patrick M Pilarski, Adam White, and Doina Precup · 2011
Earlier work this paper cites.
Learning predictive models of a depth camera & manipulator from raw execution traces
Byron Boots, Arunkumar Byravan, and Dieter Fox · 2014
Earlier work this paper cites.
Auto-encoding variational Bayes
Diederik P Kingma and Max Welling · 2014
Earlier work this paper cites.
Proximal algorithms
Neal Parikh, Stephen Boyd, et al · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Danilo J Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Greg Wayne, David Silver, Timothy Lillicrap, Yuval Tassa, and Tom Erez · 2015
Earlier work this paper cites.
DeepMPC: learning deep latent features for model predictive control
Ian Lenz, Ross Knepper, and Ashutosh Saxena · 2015
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard Lewis, and Satinder Singh · 2015
Earlier work this paper cites.
Deep learning helicopter dynamics models
Ali Punjani and Pieter Abbeel · 2015
Earlier work this paper cites.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Cited alongside, same era.
Embed to control: a locally linear latent dynamics model for control from raw images
Manuel Watter, Jost Tobias Springenberg, Joschka Boedecker, and Martin Riedmiller · 2015
Cited alongside, same era.
Deep visual foresight for planning robot motion
Chelsea Finn and Sergey Levine · 2016
Cited alongside, same era.
Deep spatial autoencoders for visuomotor learning
Chelsea Finn, Xin Yu Tan, Yan Duan, Trevor Darrell, Sergey Levine, and Pieter Abbeel · 2016
Cited alongside, same era.
Adversarial autoencoders
Alireza Makhzani, Jonathon Shlens, Navdeep Jaitly, Ian Goodfellow, and Brendan Frey · 2016
Cited alongside, same era.
Deep multi-scale video prediction beyond mean square error
Michael Mathieu, Camille Couprie, and Yann LeCun · 2016
Cited alongside, same era.
Stochastic adversarial video prediction
Alex X Lee, Richard Zhang, Frederik Ebert, Pieter Abbeel, Chelsea Finn, and Sergey Levine · 2018
Later among the works it cites.
Data-efficient hierarchical reinforcement learning
Ofir Nachum, Google Brain, Shane Gu, Honglak Lee, and Sergey Levine · 2018
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Anusha Nagabandi, Gregory Kahn, Ronald S Fearing, and Sergey Levine · 2018
Later among the works it cites.
Visual reinforcement learning with imagined goals
Ashvin Nair, Vitchyr Pong, Murtaza Dalal, Shikhar Bahl, Steven Lin, and Sergey Levine · 2018
Later among the works it cites.
Q-map: a convolutional approach for goal-oriented reinforcement learning
Fabio Pardo, Vitaly Levdik, and Petar Kormushev · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Value iteration networks
Aviv Tamar, Yi Wu, Garrett Thomas, Sergey Levine, and Pieter Abbeel · 2016
Cited alongside, same era.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob Mcgrew, Josh Tobin, Pieter Abbeel, and Wojciech Zaremba · 2017
Cited alongside, same era.
Recurrent environment simulators
Silvia Chiappa, Sébastien Racaniere, Daan Wierstra, and Shakir Mohamed · 2017
Cited alongside, same era.
Self-supervised visual planning with temporal skip connections
Frederik Ebert, Chelsea Finn, Alex X Lee, and Sergey Levine · 2017
Cited alongside, same era.
Video pixel networks
Nal Kalchbrenner, Aäron van den Oord, Karen Simonyan, Ivo Danihelka, Oriol Vinyals, Alex Graves, and Koray Kavukcuoglu · 2017
Cited alongside, same era.
Value prediction network
Junhyuk Oh, Satinder Singh, and Honglak Lee · 2017
Cited alongside, same era.
Multi-goal reinforcement learning: challenging robotics environments and request for research
Matthias Plappert, Marcin Andrychowicz, Alex Ray, Bob Mcgrew, Bowen Baker, Glenn Powell, Jonas Schneider, Josh Tobin, Maciek Chociej, Peter Welinder, Vikash Kumar, and Wojciech Zaremba · 2018
Later among the works it cites.
Temporal difference models: model-free deep rl For model-based control
Vitchyr Pong, Shixiang Gu, Murtaza Dalal, and Sergey Levine · 2018
Later among the works it cites.
Many-goals reinforcement learning
Vivek Veeriah, Junhyuk Oh, and Satinder Singh · 2018
Later among the works it cites.
Unsupervised control through non-parametric discriminative rewards
David Warde-Farley, Tom Van de Wiele, Tejas Kulkarni, Catalin Ionescu, Steven Hansen, and Volodymyr Mnih · 2018
Later among the works it cites.
Model learning for look-ahead exploration in continuous control
Arpit Agarwal, Katharina Muelling, and Katerina Fragkiadaki · 2019
Closest in time.
CURIOUS: intrinsically motivated multi-task, multi-goal reinforcement learning
Cédric Colas, Pierre Fournier, Olivier Sigaud, and Pierre-Yves Oudeyer · 2019
Closest in time.
Diagnosing and enhancing vae models
Bin Dai and David Wipf · 2019
Closest in time.
Learning actionable representations with goal-conditioned policies
Dibya Ghosh, Abhishek Gupta, and Sergey Levine · 2019
Closest in time.
An investigation of model-free planning
Arthur Guez, Mehdi Mirza, Karol Gregor, Rishabh Kabra, Sébastien Racanière, Théophane Weber, David Raposo, Adam Santoro, Laurent Orseau, Tom Eccles, et al · 2019
Closest in time.
Learning latent dynamics for planning from pixels
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson · 2019
Closest in time.
Model-based reinforcement learning for atari
Lukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski, Roy H Campbell, Konrad Czechowski, Dumitru Erhan, Chelsea Finn, Piotr Kozakowski, Sergey Levine, et al · 2019
Closest in time.
Learning multi-level hierarchies with hindsight
Andrew Levy, George Konidaris, Robert Platt, and Kate Saenko · 2019
Closest in time.
Skew-fit: state-covering self-supervised reinforcement learning
Vitchyr H. Pong, Murtaza Dalal, Steven Lin, Ashvin Nair, Shikhar Bahl, and Sergey Levine · 2019
Closest in time.
Solar: deep structured latent representations for model-based reinforcement learning
Marvin Zhang, Sharad Vikram, Laura Smith, Pieter Abbeel, Matthew J Johnson, and Sergey Levine · 2019
Closest in time.