Fetching the paper…
Reading the bibliography…
Model-based reinforcement learning could enable sample-efficient learning by quickly acquiring rich knowledge about the world and using it to improve behaviour without additional data.
An empirical Bayes approach to statistics
H. Robbins · 1956
Earlier work this paper cites.
An empirical Bayes estimator of the mean of a normal population
K. Miyasawa · 1961
Earlier work this paper cites.
Model predictive control: theory and practices - a survey
C. E. Garcia, D. M. Prett, and M. Morari · 1989
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
R. S. Sutton · 1990
Earlier work this paper cites.
A tutorial on energy-based learning
Y. LeCun, S. Chopra, and R. Hadsell · 2006
Earlier work this paper cites.
Least squares estimation without priors or supervision
M. Raphan and E. P. Simoncelli · 2011
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
P. Vincent · 2011
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
M. Deisenroth and C. E. Rasmussen · 2011
Earlier work this paper cites.
Synthesis and stabilization of complex behaviors through online trajectory optimization
Y. Tassa, T. Erez, and E. Todorov · 2012
Earlier work this paper cites.
The cross-entropy method for optimization
Z. I. Botev, D. P. Kroese, R. Y. Rubinstein, and P. L’Ecuyer · 2013
Earlier work this paper cites.
Auto-encoding Variational Bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
Guided policy search
S. Levine and V. Koltun · 2013
Earlier work this paper cites.
What regularized auto-encoders learn from the data-generating distribution
G. Alain and Y. Bengio · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel · 2016
Cited alongside, same era.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
Optimal control with learned local models: Application to dexterous manipulation
V. Kumar, E. Todorov, and S. Levine · 2016
Cited alongside, same era.
Mastering the game of Go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. Hoof, and D. Meger · 2018
Later among the works it cites.
Model-ensemble trust-region policy optimization
T. Kurutach, I. Clavera, Y. Duan, A. Tamar, and P. Abbeel · 2018
Later among the works it cites.
Model-based reinforcement learning via meta-policy optimization
I. Clavera, J. Rothfuss, J. Schulman, Y. Fujita, T. Asfour, and P. Abbeel · 2018
Later among the works it cites.
Recurrent world models facilitate policy evolution
D. Ha and J. Schmidhuber · 2018
Later among the works it cites.
Plan online, learn offline: Efficient learning and exploration via model-based control
K. Lowrey, A. Rajeswaran, S. Kakade, E. Todorov, and I. Mordatch · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Cited alongside, same era.
Prediction and control with temporal segment models
N. Mishra, P. Abbeel, and I. Mordatch · 2017
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
K. Chua, R. Calandra, R. McAllister, and S. Levine · 2018
Cited alongside, same era.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
A. Nagabandi, G. Kahn, R. S. Fearing, and S. Levine · 2018
Cited alongside, same era.
Deep energy estimator networks
S. Saremi, A. Mehrjou, B. Schölkopf, and A. Hyvärinen · 2018
Cited alongside, same era.
PIPPS: Flexible model-based policy search robust to the curse of chaos
P. Parmas, C. E. Rasmussen, J. Peters, and K. Doya · 2018
Cited alongside, same era.
Closest in time.
Learning to adapt in dynamic, real-world environments through meta-reinforcement learning
I. Clavera, A. Nagabandi, S. Liu, R. S. Fearing, P. Abbeel, S. Levine, and C. Finn · 2019
Closest in time.
S. Saremi and A. Hyvärinen · 2019
Closest in time.
Regularizing trajectory optimization with denoising autoencoders
R. Boney, N. Di Palo, M. Berglund, A. Ilin, J. Kannala, A. Rasmus, and H. Valpola · 2019
Closest in time.
Learning latent dynamics for planning from pixels
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson · 2019
Closest in time.
Implicit generation and generalization in energy-based models
Y. Du and I. Mordatch · 2019
Closest in time.
Model based planning with energy based models
Y. Du, T. Lin, and I. Mordatch · 2019
Closest in time.