Fetching the paper…
Reading the bibliography…
In the present paper, we propose an extension of the Deep Planning Network (PlaNet), also referred to as PlaNet of the Bayesians (PlaNet-Bayes).
F. D. Vargas-Villamil and D. E. Rivera, “Multilayer optimization and scheduling using model predictive control: application to reentrant semiconductor manufacturing lines,” Computers & Chemical Engineering
2000
Earlier work this paper cites.
ACM New York, NY, USA, 2002
S. Thrun, Probabilistic robotics · 2002
Earlier work this paper cites.
N. Hansen, S. D. Müller, and P. Koumoutsakos, “Reducing the time complexity of the derandomized evolution strategy with covariance matrix adaptation (CMA-ES),” Evolutionary computation
2003
Earlier work this paper cites.
M. Deisenroth and C. E. Rasmussen, “PILCO: A model-based and data-efficient approach to policy search,” in International Conference on Machine Learning (ICML)
2011
Earlier work this paper cites.
Z. I. Botev, D. P. Kroese, R. Y. Rubinstein, and P. L’Ecuyer, “The cross-entropy method for optimization,” in Handbook of statistics
2013
Earlier work this paper cites.
S. Goschin, A. Weinstein, and M. Littman, “The cross-entropy method optimizes for quantiles,” in International Conference on Machine Learning (ICML)
2013
Earlier work this paper cites.
A. Afram and F. Janabi-Sharifi, “Theory and applications of HVAC control systems–a review of model predictive control (MPC),” Building and Environment
2014
Earlier work this paper cites.
S. Vazquez, J. Leon, L. Franquelo, J. Rodriguez, H. A. Young, A. Marquez, and P. Zanchetta, “Model predictive control: A review of its applications in power electronics,” IEEE Ind. Electron. Mag
2014
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” in International Conference on Learning Representations (ICLR)
2014
Earlier work this paper cites.
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, et al
2014
Earlier work this paper cites.
Now Publishers, Inc., 2015
S. Bubeck et al · 2015
Earlier work this paper cites.
Software available from tensorflow.org
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, et al · 2015
Earlier work this paper cites.
Y. Gal and Z. Ghahramani, “Dropout as a bayesian approximation: Representing model uncertainty in deep learning,” in International Conference on Machine Learning (ICML)
2016
Earlier work this paper cites.
G. Williams, P. Drews, B. Goldfain, J. M. Rehg, and E. A. Theodorou, “Aggressive driving with model predictive path integral control,” in International Conference on Robotics and Automation (ICRA)
2016
Earlier work this paper cites.
D. Silver, A. Huang, C. J. Maddison, A. Guez, et al
2016
Cited alongside, same era.
Y. Gal, J. Hron, and A. Kendall, “Concrete dropout,” in Neural Information Processing Systems
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
K. Chua, R. Calandra, R. McAllister, and S. Levine, “Deep reinforcement learning in a handful of trials using probabilistic dynamics models,” in Neural Information Processing Systems
2018
Cited alongside, same era.
V. Thomas, E. Bengio, W. Fedus, J. Pondard, et al
2018
Later among the works it cites.
Y. Sawada, “Disentangling controllable and uncontrollable factors of variation by interacting with the world,” in NeurIPS Deep Reinforcement Learning Workshop
2018
Later among the works it cites.
W. H. Beluch, T. Genewein, A. Nürnberger, and J. M. Köhler, “The power of ensembles for active learning in image classification,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2018
Later among the works it cites.
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson, “Learning latent dynamics for planning from pixels,” in International Conference on Machine Learning (ICML)
2019
Later among the works it cites.
M. Okada and T. Taniguchi, “Variational inference MPC for bayesian model-based reinforcement learning,” in Conference on Robot Learning (CoRL)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Kurutach, I. Clavera, Y. Duan, A. Tamar, and P. Abbeel, “Model-ensemble trust-region policy optimization,” in International Conference on Learning Representations (ICLR)
2018
Cited alongside, same era.
I. Clavera, J. Rothfuss, J. Schulman, Y. Fujita, T. Asfour, and P. Abbeel, “Model-based reinforcement learning via meta-policy optimization,” in Conference on Robot Learning (CoRL)
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, et al
2018
Cited alongside, same era.
M. Okada and T. Taniguchi, “Acceleration of gradient-based path integral method for efficient optimal and inverse optimal control,” in International Conference on Robotics and Automation (ICRA)
2018
Cited alongside, same era.
D. Ha and J. Schmidhuber, “World models,” arXiv preprint arXiv:1803.10122
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Later among the works it cites.
A. Nagabandi, K. Konoglie, S. Levine, and V. Kumar, “Deep dynamics models for learning dexterous manipulation,” in Conference on Robot Learning (CoRL)
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
L. Kaiser, M. Babaeizadeh, P. Milos, B. Osinski, et al
2019
Later among the works it cites.
D. Tran, M. Dusenberry, M. van der Wilk, and D. Hafner, “Bayesian layers: A module for neural network uncertainty,” in Neural Information Processing Systems
2019
Later among the works it cites.
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, et al
2019
Later among the works it cites.
D. Han, K. Doya, and J. Tani, “Variational recurrent models for solving partially observable control tasks,” in International Conference on Learning Representations (ICLR)
2020
Closest in time.