Fetching the paper…
Reading the bibliography…
By planning through a learned dynamics model, model-based reinforcement learning (MBRL) offers the prospect of good performance with little environment interaction.
Markov decision processes
Puterman, M. L · 1990
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, R. S · 1990
Earlier work this paper cites.
Mixture density networks
Bishop, C. M · 1994
Earlier work this paper cites.
Solving uncertain markov decision processes
Bagnell, J. A., Ng, A. Y., and Schneider, J. G · 2001
Earlier work this paper cites.
Iterative linear quadratic regulator design for nonlinear biological movement systems
Li, W. and Todorov, E · 2004
Earlier work this paper cites.
Robust solutions to markov decision problems with uncertain transition matrices
El Ghaoui, L. and Nilim, A · 2005
Earlier work this paper cites.
Using inaccurate models in reinforcement learning
Abbeel, P., Quigley, M., and Ng, A. Y · 2006
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Deisenroth, M. and Rasmussen, C. E · 2011
Earlier work this paper cites.
Discovery of complex behaviors through contact-invariant optimization
Mordatch, I., Todorov, E., and Popović, Z · 2012
Earlier work this paper cites.
Density ratio estimation in machine learning
Sugiyama, M., Suzuki, T., and Kanamori, T · 2012
Earlier work this paper cites.
Synthesis and stabilization of complex behaviors through online trajectory optimization
Tassa, Y., Erez, T., and Todorov, E · 2012
Earlier work this paper cites.
Gaussian processes for data-efficient learning in robotics and control
Deisenroth, M. P., Fox, D., and Rasmussen, C. E · 2013
Earlier work this paper cites.
The cross-entropy method: a unified approach to combinatorial optimization, Monte-Carlo simulation and machine learning
Rubinstein, R. Y. and Kroese, D. P · 2013
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
Heess, N., Wayne, G., Silver, D., Lillicrap, T., Erez, T., and Tassa, Y · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Model predictive path integral control using covariance variable importance sampling
Williams, G., Aldrich, A., and Theodorou, E · 2015
Cited alongside, same era.
Improving pilco with bayesian neural network dynamics models
Gal, Y., McAllister, R., and Rasmussen, C. E · 2016
Cited alongside, same era.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Nagabandi, A., Kahn, G., Fearing, R. S., and Levine, S · 2018
Later among the works it cites.
Bias correction of learned generative models using likelihood-free importance weighting
Grover, A., Song, J., Kapoor, A., Tran, K., Agarwal, A., Horvitz, E. J., and Ermon, S · 2019
Later among the works it cites.
Learning latent dynamics for planning from pixels
Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., and Davidson, J · 2019
Later among the works it cites.
When to trust your model: Model-based policy optimization
Janner, M., Fu, J., Zhang, M., and Levine, S · 2019
Later among the works it cites.
Model-based reinforcement learning for atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimal control with learned local models: Application to dexterous manipulation
Kumar, V., Todorov, E., and Levine, S · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P · 2016
Cited alongside, same era.
Hierarchical implicit models and likelihood-free variational inference
Tran, D., Ranganath, R., and Blei, D · 2017
Cited alongside, same era.
Do gans learn the distribution? some theory and empirics
Arora, S., Risteski, A., and Zhang, Y · 2018
Cited alongside, same era.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Buckman, J., Hafner, D., Tucker, G., Brevdo, E., and Lee, H · 2018
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S · 2018
Cited alongside, same era.
Model-based value estimation for efficient model-free reinforcement learning
Feinberg, V., Wan, A., Stoica, I., Jordan, M. I., Gonzalez, J. E., and Levine, S · 2018
Cited alongside, same era.
Langlois, E., Zhang, S., Zhang, G., Abbeel, P., and Ba, J · 2019
Later among the works it cites.
Dynamics-aware unsupervised discovery of skills
Sharma, A., Gu, S., Levine, S., Kumar, V., and Hausman, K · 2019
Later among the works it cites.
Exploring model-based planning with policy networks
Wang, T. and Ba, J · 2019
Later among the works it cites.
Off-dynamics reinforcement learning: Training for transfer with domain classifiers
Eysenbach, B., Asawa, S., Chaudhari, S., Salakhutdinov, R., and Levine, S · 2020
Later among the works it cites.
Objective mismatch in model-based reinforcement learning
Lambert, N., Amos, B., Yadan, O., and Calandra, R · 2020
Later among the works it cites.
Deep dynamics models for learning dexterous manipulation
Nagabandi, A., Konolige, K., Levine, S., and Kumar, V · 2020
Later among the works it cites.
Goal-aware prediction: Learning to model what matters
Nair, S., Savarese, S., and Finn, C · 2020
Later among the works it cites.