Fetching the paper…
Reading the bibliography…
Model-based reinforcement learning (RL) algorithms can attain excellent sample efficiency, but often lag behind the best model-free algorithms in terms of asymptotic performance.
A Markovian decision process
R. Bellman · 1957
Earlier work this paper cites.
A discussion of random methods for seeking maxima
S. H. Brooks · 1958
Earlier work this paper cites.
Neural network modeling and an extended DMC algorithm to control nonlinear systems
E. Hernandaz and Y. Arkun · 1990
Earlier work this paper cites.
Real-time dynamic control of an industrial manipulator using a neural network-based learning controller
W. T. Miller, R. P. Hewes, F. H. Glanz, and L. G. Kraft · 1990
Earlier work this paper cites.
Reinforcement Learning for Robots Using Neural Networks
L.-J. Lin · 1992
Earlier work this paper cites.
A practical Bayesian framework for backpropagation networks
D. J. MacKay · 1992
Earlier work this paper cites.
Efficient exploration in reinforcement learning
S. Thrun · 1992
Earlier work this paper cites.
An introduction to the bootstrap
B. Efron and R. Tibshirani · 1994
Earlier work this paper cites.
Model predictive control using neural networks
A. Draeger, S. Engell, and H. Ranke · 1995
Earlier work this paper cites.
Bayesian learning for neural networks
R. Neal · 1995
Earlier work this paper cites.
A comparison of direct and model-based reinforcement learning
C. G. Atkeson and J. C. Santamaría · 1997
Earlier work this paper cites.
Multiple-step ahead prediction for non linear dynamic systems–a Gaussian process treatment with propagation of the uncertainty
A. Girard, C. E. Rasmussen, J. Quinonero-Candela, R. Murray-Smith, O. Winther, and J. Larsen · 2002
Earlier work this paper cites.
Propagation of uncertainty in Bayesian kernel models—application to multiple-step ahead forecasting
J. Quiñonero-Candela, A. Girard, J. Larsen, and C. E. Rasmussen · 2003
Earlier work this paper cites.
Gaussian processes in reinforcement learning
C. E. Rasmussen and M. Kuss · 2003
Earlier work this paper cites.
Gaussian process model based predictive control
J. Kocijan, R. Murray-Smith, C. E. Rasmussen, and A. Girard · 2004
Earlier work this paper cites.
Using inaccurate models in reinforcement learning
P. Abbeel, M. Quigley, and A. Y. Ng · 2006
Earlier work this paper cites.
Gaussian processes and reinforcement learning for identification and control of an autonomous blimp
J. Ko, D. J. Klein, D. Fox, and D. Haehnel · 2007
Earlier work this paper cites.
Explicit stochastic predictive control of combustion plants based on Gaussian process models
A. Grancharova, J. Kocijan, and T. A. Johansen · 2008
Earlier work this paper cites.
Local Gaussian process regression for real time online model learning
D. Nguyen-Tuong, J. Peters, and M. Seeger · 2008
Earlier work this paper cites.
Policy search for motor primitives in robotics
J. Kober and J. Peters · 2009
Earlier work this paper cites.
Active learning of inverse models with intrinsically motivated goal exploration in robots
A. Baranes and P.-Y. Oudeyer · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Cited alongside, same era.
The cross-entropy method for optimization
Z. I. Botev, D. P. Kroese, R. Y. Rubinstein, and P. L’Ecuyer · 2013
Cited alongside, same era.
Model predictive control
E. F. Camacho and C. B. Alba · 2013
Cited alongside, same era.
Gaussian processes for data-efficient learning in robotics and control
M. Deisenroth, D. Fox, and C. Rasmussen · 2013
Cited alongside, same era.
Data-efficient generalization of robot skills with contextual policy search
A. G. Kupcsik, M. P. Deisenroth, J. Peters, and G. Neumann · 2013
Cited alongside, same era.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Later among the works it cites.
Combining model-based policy search with online model learning for control of physical humanoids
I. Mordatch, N. Mishra, C. Eppner, and P. Abbeel · 2016
Later among the works it cites.
Risk versus uncertainty in deep learning: Bayes, bootstrap and the dangers of dropout
I. Osband · 2016
Later among the works it cites.
Deep exploration via bootstrapped DQN
I. Osband, C. Blundell, A. Pritzel, and B. Van Roy · 2016
Later among the works it cites.
Combining model-based and model-free updates for trajectory-centric reinforcement learning
Y. Chebotar, K. Hausman, M. Zhang, G. Sukhatme, S. Schaal, and S. Levine · 2017
Later among the works it cites.
Openai baselines
P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, and Y. Wu · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Weight uncertainty in neural networks
C. Blundell, J. Cornebise, K. Kavukcuoglu, and D. Wierstra · 2015
Cited alongside, same era.
Probabilistic backpropagation for scalable learning of Bayesian neural networks
J. M. Hernández-Lobato and R. Adams · 2015
Cited alongside, same era.
DeepMPC: Learning deep latent features for model predictive control
I. Lenz, R. Knepper, and A. Saxena · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Deep learning helicopter dynamics models
A. Punjani and P. Abbeel · 2015
Cited alongside, same era.
Learning to poke by poking: Experiential learning of intuitive physics
P. Agrawal, A. V. Nair, P. Abbeel, J. Malik, and S. Levine · 2016
Cited alongside, same era.
Later among the works it cites.
Concrete dropout
Y. Gal, J. Hron, and A. Kendall · 2017
Later among the works it cites.
On calibration of modern neural networks
C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger · 2017
Later among the works it cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
B. Lakshminarayanan, A. Pritzel, and C. Blundell · 2017
Later among the works it cites.
Data-efficient reinforcement learning in continuous state-action Gaussian-POMDPs
R. McAllister and C. E. Rasmussen · 2017
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
A. Nagabandi, G. Kahn, R. S. Fearing, and S. Levine · 2017
Later among the works it cites.
Searching for activation functions
P. Ramachandran, B. Zoph, and Q. V. Le · 2017
Later among the works it cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Later among the works it cites.
Information theoretic MPC for model-based reinforcement learning
G. Williams, N. Wagener, B. Goldfain, P. Drews, J. M. Rehg, B. Boots, and E. A. Theodorou · 2017
Later among the works it cites.
Decomposition of uncertainty in Bayesian deep learning for efficient and risk-sensitive learning
S. Depeweg, J.-M. Hernandez-Lobato, F. Doshi-Velez, and S. Udluft · 2018
Closest in time.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Closest in time.
Synthesizing neural network controllers with probabilistic model based reinforcement learning
J. C. G. Higuera, D. Meger, and G. Dudek · 2018
Closest in time.
Data-efficient reinforcement learning with probabilistic model predictive control
S. Kamthe and M. P. Deisenroth · 2018
Closest in time.
Model-ensemble trust-region policy optimization
T. Kurutach, I. Clavera, Y. Duan, A. Tamar, and P. Abbeel · 2018
Closest in time.
PIPPS: Flexible model-based policy search robust to the curse of chaos
P. Parmas, C. E. Rasmussen, J. Peters, and K. Doya · 2018
Closest in time.