Fetching the paper…
Reading the bibliography…
In recent studies on model-based reinforcement learning (MBRL), incorporating uncertainty in forward dynamics is a state-of-the-art strategy to enhance learning performance, making MBRLs competitive to cutting-edge model free methods, especially in simulated robotics tasks.
A gentle tutorial of the EM algorithm and its application to parameter estimation for gaussian mixture and hidden markov models
J. A. Bilmes et al · 1998
Earlier work this paper cites.
Multilayer optimization and scheduling using model predictive control: application to reentrant semiconductor manufacturing lines
F. D. Vargas-Villamil and D. E. Rivera · 2000
Earlier work this paper cites.
The infinite gaussian mixture model
C. E. Rasmussen · 2000
Earlier work this paper cites.
Reducing the time complexity of the derandomized evolution strategy with covariance matrix adaptation (CMA-ES)
N. Hansen, S. D. Müller, and P. Koumoutsakos · 2003
Earlier work this paper cites.
Sequential monte carlo for model predictive control
N. Kantas, J. Maciejowski, and A. Lecchini-Visintini · 2009
Earlier work this paper cites.
A generalized path integral control approach to reinforcement learning
E. Theodorou, J. Buchli, and S. Schaal · 2010
Earlier work this paper cites.
PILCO: A model-based and data-efficient approach to policy search
M. Deisenroth and C. E. Rasmussen · 2011
Earlier work this paper cites.
Handbook of Markov Chain Monte Carlo , chapter 11
S. Brooks, A. Gelman, G. Jones, and X.-L. Meng · 2011
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
The cross-entropy method for optimization
Z. I. Botev, D. P. Kroese, R. Y. Rubinstein, and P. L’Ecuyer · 2013
Earlier work this paper cites.
The cross-entropy method optimizes for quantiles
S. Goschin, A. Weinstein, and M. Littman · 2013
Earlier work this paper cites.
Theory and applications of HVAC control systems–a review of model predictive control (MPC)
A. Afram and F. Janabi-Sharifi · 2014
Earlier work this paper cites.
Model predictive control: A review of its applications in power electronics
S. Vazquez, J. Leon, L. Franquelo, J. Rodriguez, H. A. Young, A. Marquez, and P. Zanchetta · 2014
Earlier work this paper cites.
Convex optimization: Algorithms and complexity , volume 8, chapter 4
S. Bubeck et al · 2015
Earlier work this paper cites.
Model-based relative entropy stochastic search
A. Abdolmaleki, R. Lioutikov, J. R. Peters, N. Lau, L. P. Reis, and G. Neumann · 2015
Cited alongside, same era.
A survey of motion planning and control techniques for self-driving urban vehicles
B. Paden, M. Čáp, S. Z. Yong, D. Yershov, and E. Frazzoli · 2016
Cited alongside, same era.
Optimization-based locomotion planning, estimation, and control design for the Atlas humanoid robot
S. Kuindersma, R. Deits, M. Fallon, A. Valenzuela, H. Dai, F. Permenter, T. Koolen, P. Marion, and R. Tedrake · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Cited alongside, same era.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Y. Gal and Z. Ghahramani · 2016
Cited alongside, same era.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
A. Nagabandi, G. Kahn, R. S. Fearing, and S. Levine · 2018
Later among the works it cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
K. Chua, R. Calandra, R. McAllister, and S. Levine · 2018
Later among the works it cites.
Model-ensemble trust-region policy optimization
T. Kurutach, I. Clavera, Y. Duan, A. Tamar, and P. Abbeel · 2018
Later among the works it cites.
Model-based reinforcement learning via meta-policy optimization
I. Clavera, J. Rothfuss, J. Schulman, Y. Fujita, T. Asfour, and P. Abbeel · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aggressive driving with model predictive path integral control
G. Williams, P. Drews, B. Goldfain, J. M. Rehg, and E. A. Theodorou · 2016
Cited alongside, same era.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Cited alongside, same era.
Information theoretic MPC for model-based reinforcement learning
G. Williams, N. Wagener, B. Goldfain, P. Drews, J. M. Rehg, B. Boots, and E. A. Theodorou · 2017
Cited alongside, same era.
Concrete dropout
Y. Gal, J. Hron, and A. Kendall · 2017
Cited alongside, same era.
Uncertainty-aware reinforcement learning for collision avoidance
G. Kahn, A. Villaflor, V. Pong, P. Abbeel, and S. Levine · 2017
Cited alongside, same era.
Dropout inference in bayesian neural networks with alpha-divergences
Y. Li and Y. Gal · 2017
Cited alongside, same era.
Deriving and improving CMA-ES with information geometric trust regions
A. Abdolmaleki, B. Price, N. Lau, L. P. Reis, and G. Neumann · 2017
Cited alongside, same era.
S. Levine · 2018
Later among the works it cites.
Mirror descent search and its acceleration
M. Miyashita, S. Yano, and T. Kondo · 2018
Later among the works it cites.
Acceleration of gradient-based path integral method for efficient optimal and inverse optimal control
M. Okada and T. Taniguchi · 2018
Later among the works it cites.
Probabilistic planning with sequential monte carlo methods
A. Piche, V. Thomas, C. Ibrahim, Y. Bengio, and C. Pal · 2018
Later among the works it cites.
A bayesian approach to generative adversarial imitation learning
W. Jeon, S. Seo, and K.-E. Kim · 2018
Later among the works it cites.
An online learning approach to model predictive control
N. Wagener, C.-A. Cheng, J. Sacks, and B. Boots · 2019
Closest in time.
A. Kinose and T. Tadahiro · 2019
Closest in time.
Learning latent dynamics for planning from pixels
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson · 2019
Closest in time.