Fetching the paper…
Reading the bibliography…
Model-based reinforcement learning is a compelling framework for data-efficient learning of agents that interact with the world.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala · 1912
Earlier work this paper cites.
A markovian decision process
R. Bellman · 1957
Earlier work this paper cites.
An introduction to the bootstrap
B. Efron and R. J. Tibshirani · 1994
Earlier work this paper cites.
The elements of statistical learning , volume 1
J. Friedman, T. Hastie, and R. Tibshirani · 2001
Earlier work this paper cites.
A tutorial on the cross-entropy method
P.-T. De Boer, D. P. Kroese, S. Mannor, and R. Y. Rubinstein · 2005
Earlier work this paper cites.
The numpy array: a structure for efficient numerical computation
S. Van Der Walt, S. C. Colbert, and G. Varoquaux · 2011
Earlier work this paper cites.
Model predictive control
E. F. Camacho and C. B. Alba · 2013
Earlier work this paper cites.
Gaussian processes for data-efficient learning in robotics and control
M. P. Deisenroth, D. Fox, and C. E. Rasmussen · 2013
Earlier work this paper cites.
Guided policy search
S. Levine and V. Koltun · 2013
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
N. Heess, G. Wayne, D. Silver, T. Lillicrap, T. Erez, and Y. Tassa · 2015
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, et al · 2016
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Continuous deep q-learning with model-based acceleration
S. Gu, T. Lillicrap, I. Sutskever, and S. Levine · 2016
Earlier work this paper cites.
Openai baselines
P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, Y. Wu, and P. Zhokhov · 2017
Earlier work this paper cites.
Learning model-based planning from scratch
R. Pascanu, Y. Li, O. Vinyals, N. Heess, L. Buesing, S. Racanière, D. Reichert, T. Weber, D. Wierstra, and P. Battaglia · 2017
Earlier work this paper cites.
Imagination-augmented agents for deep reinforcement learning
T. Weber, S. Racanière, D. P. Reichert, L. Buesing, A. Guez, D. J. Rezende, A. P. Badia, O. Vinyals, N. Heess, Y. Li, et al · 2017
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
K. Chua, R. Calandra, R. McAllister, and S. Levine · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Model-predictive policy learning with uncertainty regularization for driving in dense traffic
M. Henaff, A. Canziani, and Y. LeCun · 2018
Cited alongside, same era.
Deep reinforcement learning that matters
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger · 2018
Cited alongside, same era.
The gap between model-based and model-free methods on the linear quadratic regulator: An asymptotic viewpoint
S. Tu and B. Recht · 2019
Later among the works it cites.
Exploring model-based planning with policy networks
T. Wang and J. Ba · 2019
Later among the works it cites.
Benchmarking model-based reinforcement learning
T. Wang, X. Bao, I. Clavera, J. Hoang, Y. Wen, E. Langlois, S. Zhang, G. Zhang, P. Abbeel, and J. Ba · 2019
Later among the works it cites.
Hydra - a framework for elegantly configuring complex applications
O. Yadan · 2019
Later among the works it cites.
The differentiable cross-entropy method
B. Amos and D. Yarats · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pytorch implementations of reinforcement learning algorithms
I. Kostrikov · 2018
Cited alongside, same era.
Model-ensemble trust-region policy optimization
T. Kurutach, I. Clavera, Y. Duan, A. Tamar, and P. Abbeel · 2018
Cited alongside, same era.
Plan online, learn offline: Efficient learning and exploration via model-based control
K. Lowrey, A. Rajeswaran, S. Kakade, E. Todorov, and I. Mordatch · 2018
Cited alongside, same era.
Imagined value gradients: Model-based policy optimization with transferable latent dynamics models
A. Byravan, J. T. Springenberg, A. Abdolmaleki, R. Hafner, M. Neunert, T. Lampe, N. Siegel, N. Heess, and M. Riedmiller · 2019
Cited alongside, same era.
TF-Agents: A library for reinforcement learning in tensorflow
S. Guadarrama, A. Korattikara, O. Ramirez, P. Castro, E. Holly, S. Fishman, K. Wang, E. Gonina, N. Wu, E. Kokiopoulou, L. Sbaiz, J. Smith, G. Bartók, J. Berent, C. Harris, V. Vanhoucke, and E. Brevdo · 2019
Cited alongside, same era.
Learning latent dynamics for planning from pixels
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson · 2019
Cited alongside, same era.
When to trust your model: Model-based policy optimization
M. Janner, J. Fu, M. Zhang, and S. Levine · 2019
Cited alongside, same era.
J. B. Hamrick, A. L. Friesen, F. Behbahani, A. Guez, F. Viola, S. Witherspoon, T. Anthony, L. Buesing, P. Veličković, and T. Weber · 2020
Later among the works it cites.
Objective mismatch in model-based reinforcement learning
N. O. Lambert, B. Amos, O. Yadan, and R. Calandra · 2020
Later among the works it cites.
Deep dynamics models for learning dexterous manipulation
A. Nagabandi, K. Konolige, S. Levine, and V. Kumar · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, et al · 2020
Later among the works it cites.
dm_control: Software and tasks for continuous control, 2020
Y. Tassa, S. Tunyasuvunakool, A. Muldal, Y. Doron, S. Liu, S. Bohez, J. Merel, T. Erez, T. Lillicrap, and N. Heess · 2020
Later among the works it cites.
Soft actor-critic (sac) implementation in pytorch
D. Yarats and I. Kostrikov · 2020
Later among the works it cites.
On the stochastic value gradient for continuous reinforcement learning
B. Amos, S. Stanton, D. Yarats, and A. G. Wilson · 2021
Closest in time.
Bellman: A toolbox for model-based reinforcement learning in tensorflow
J. McLeod, H. Stojic, V. Adam, D. Kim, J. Grau-Moya, P. Vrancx, and F. Leibfried · 2021
Closest in time.
On the importance of hyperparameter optimization for model-based reinforcement learning
B. Zhang, R. Rajan, L. Pineda, N. Lambert, A. Biedenkapp, K. Chua, F. Hutter, and R. Calandra · 2021
Closest in time.