Fetching the paper…
Reading the bibliography…
Model-based reinforcement learning approaches leverage a forward dynamics model to support planning and decision making, which, however, may fail catastrophically if the model is inaccurate.
Dyna, an integrated architecture for learning, planning, and reacting
Sutton, R. S · 1991
Earlier work this paper cites.
Model predictive control using neural networks
Draeger, A., Engell, S., and Ranke, H · 1995
Earlier work this paper cites.
Gaussian processes in reinforcement learning
Kuss, M. and Rasmussen, C. E · 2004
Earlier work this paper cites.
Gaussian processes and reinforcement learning for identification and control of an autonomous blimp
Ko, J., Klein, D. J., Fox, D., and Haehnel, D · 2007
Earlier work this paper cites.
An analysis of linear models, linear value-function approximation, and feature selection for reinforcement learning
Parr, R., Li, L., Taylor, G., Painter-Wakefield, C., and Littman, M. L · 2008
Earlier work this paper cites.
Model-based reinforcement learning with nearly tight exploration complexity bounds
Szita, I. and Szepesvári, C · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Deisenroth, M. and Rasmussen, C. E · 2011
Earlier work this paper cites.
Dyna-style planning with linear function approximation and prioritized sweeping
Sutton, R. S., Szepesvári, C., Geramifard, A., and Bowling, M. P · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Model predictive control
Camacho, E. F. and Alba, C. B · 2013
Earlier work this paper cites.
A survey on policy search for robotics
Deisenroth, M. P., Neumann, G., Peters, J., et al · 2013
Earlier work this paper cites.
Guided policy search
Levine, S. and Koltun, V · 2013
Earlier work this paper cites.
Learning neural network policies with guided policy search under unknown dynamics
Levine, S. and Abbeel, P · 2014
Earlier work this paper cites.
Conditional generative adversarial nets
Mirza, M. and Osindero, S · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Improving multi-step prediction of learned time series models
Venkatraman, A., Hebert, M., and Bagnell, J. A · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Improving pilco with bayesian neural network dynamics models
Gal, Y., McAllister, R., and Rasmussen, C. E · 2016
Cited alongside, same era.
Nips 2016 tutorial: Generative adversarial networks
Goodfellow, I · 2016
Cited alongside, same era.
Model-ensemble trust-region policy optimization
Kurutach, T., Clavera, I., Duan, Y., Tamar, A., and Abbeel, P · 2018
Later among the works it cites.
Algorithmic framework for model-based deep reinforcement learning with theoretical guarantees
Luo, Y., Xu, H., Li, Y., Tian, Y., Darrell, T., and Ma, T · 2018
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Nagabandi, A., Kahn, G., Fearing, R. S., and Levine, S · 2018
Later among the works it cites.
Dual policy iteration
Sun, W., Gordon, G. J., Boots, B., and Bagnell, J · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimal control with learned local models: Application to dexterous manipulation
Kumar, V., Todorov, E., and Levine, S · 2016
Cited alongside, same era.
Epopt: Learning robust neural network policies using model ensembles
Rajeswaran, A., Ghotra, S., Ravindran, B., and Levine, S · 2016
Cited alongside, same era.
On the sample complexity of the linear quadratic regulator
Dean, S., Mania, H., Matni, N., Recht, B., and Tu, S · 2017
Cited alongside, same era.
Prediction and control with temporal segment models
Mishra, N., Abbeel, P., and Mordatch, I · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Self-correcting models for model-based reinforcement learning
Talvitie, E · 2017
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S · 2018
Cited alongside, same era.
Understanding the asymptotic performance of model-based rl methods
Whitney, W. and Fergus, R · 2018
Later among the works it cites.
Combating the compounding-error problem with a multi-step model
Asadi, K., Misra, D., Kim, S., and Littman, M. L · 2019
Later among the works it cites.
Learning to predict without looking ahead: World models without forward prediction
Freeman, C. D., Metz, L., and Ha, D · 2019
Later among the works it cites.
When to trust your model: Model-based policy optimization
Janner, M., Fu, J., Zhang, M., and Levine, S · 2019
Later among the works it cites.
Model-based reinforcement learning for atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al · 2019
Later among the works it cites.
Benchmarking model-based reinforcement learning
Langlois, E., Zhang, S., Zhang, G., Abbeel, P., and Ba, J · 2019
Later among the works it cites.
Exploring model-based planning with policy networks
Wang, T. and Ba, J · 2019
Later among the works it cites.
Model imitation for model-based reinforcement learning
Wu, Y.-H., Fan, T.-H., Ramadge, P. J., and Su, H · 2019
Later among the works it cites.
Learning to combat compounding-error in model-based reinforcement learning
Xiao, C., Wu, Y., Ma, C., Schuurmans, D., and Müller, M · 2019
Later among the works it cites.