Fetching the paper…
Reading the bibliography…
Designing effective model-based reinforcement learning algorithms is difficult because the ease of data generation must be weighed against the bias of model-generated data.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, R. S · 1990
Earlier work this paper cites.
Model predictive control using neural networks
Draeger, A., Engell, S., and Ranke, H · 1995
Earlier work this paper cites.
Reinforcement learning: A survey
Kaelbling, L. P., Littman, M. L., and Moore, A. P · 1996
Earlier work this paper cites.
Learning tasks from a single demonstration
Atkeson, C. G. and Schaal, S · 1997
Earlier work this paper cites.
Model-based reinforcement learning with nearly tight exploration complexity bounds
Szita, I. and Szepesvari, C · 2010
Earlier work this paper cites.
PILCO: A model-based and data-efficient approach to policy search
Deisenroth, M. and Rasmussen, C. E · 2011
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Guided policy search
Levine, S. and Koltun, V · 2013
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S. and Ben-David, S · 2014
Earlier work this paper cites.
Model regularization for stable sample rollouts
Talvitie, E · 2014
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
Heess, N., Wayne, G., Silver, D., Lillicrap, T., Tassa, Y., and Erez, T · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Earlier work this paper cites.
Action-conditional video prediction using deep networks in Atari games
Oh, J., Guo, X., Lee, H., Lewis, R., and Singh, S · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
Learning and policy search in stochastic dynamical systems with bayesian neural networks
Depeweg, S., Hernández-Lobato, J. M., Doshi-Velez, F., and Udluft, S · 2016
Earlier work this paper cites.
Improving PILCO with Bayesian neural network dynamics models
Gal, Y., McAllister, R., and Rasmussen, C. E · 2016
Cited alongside, same era.
Continuous deep Q-learning with model-based acceleration
Gu, S., Lillicrap, T., Sutskever, I., and Levine, S · 2016
Cited alongside, same era.
Optimal control with learned local models: Application to dexterous manipulation
Kumar, V., Todorov, E., and Levine, S · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Cited alongside, same era.
Self-correcting models for model-based reinforcement learning
Talvitie, E · 2016
Cited alongside, same era.
Value iteration networks
Tamar, A., WU, Y., Thomas, G., Levine, S., and Abbeel, P · 2016
Cited alongside, same era.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Buckman, J., Hafner, D., Tucker, G., Brevdo, E., and Lee, H · 2018
Later among the works it cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S · 2018
Later among the works it cites.
Model-based reinforcement learning via meta-policy optimization
Clavera, I., Rothfuss, J., Schulman, J., Fujita, Y., Asfour, T., and Abbeel, P · 2018
Later among the works it cites.
Visual foresight: Model-based deep reinforcement learning for vision-based robotic control
Ebert, F., Finn, C., Dasari, S., Xie, A., Lee, A. X., and Levine, S · 2018
Later among the works it cites.
Model-based value estimation for efficient model-free reinforcement learning
Feinberg, V., Wan, A., Stoica, I., Jordan, M. I., Gonzalez, J. E., and Levine, S · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the sample complexity of the linear quadratic regulator
Dean, S., Mania, H., Matni, N., Recht, B., and Tu, S · 2017
Cited alongside, same era.
Value-aware loss function for model-based reinforcement learning
Farahmand, A.-M., Barreto, A., and Nikovski, D · 2017
Cited alongside, same era.
Uncertainty-driven imagination for continuous deep reinforcement learning
Kalweit, G. and Boedecker, J · 2017
Cited alongside, same era.
Value prediction network
Oh, J., Singh, S., and Lee, H · 2017
Cited alongside, same era.
Imagination-augmented agents for deep reinforcement learning
Racanière, S., Weber, T., Reichert, D., Buesing, L., Guez, A., Jimenez Rezende, D., Puigdomènech Badia, A., Vinyals, O., Heess, N., Li, Y., Pascanu, R., Battaglia, P., Hassabis, D., Silver, D., and Wierstra, D · 2017
Cited alongside, same era.
EPOpt: Learning robust neural network policies using model ensembles
Rajeswaran, A., Ghotra, S., Levine, S., and Ravindran, B · 2017
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Later among the works it cites.
The effect of planning shape on dyna-style planning in high-dimensional state spaces
Holland, G. Z., Talvitie, E. J., and Bowling, M · 2018
Later among the works it cites.
Model-ensemble trust-region policy optimization
Kurutach, T., Clavera, I., Duan, Y., Tamar, A., and Abbeel, P · 2018
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Nagabandi, A., Kahn, G., S. Fearing, R., and Levine, S · 2018
Later among the works it cites.
Dual policy iteration
Sun, W., Gordon, G. J., Boots, B., and Bagnell, J · 2018
Later among the works it cites.
Understanding the asymptotic performance of model-based RL methods
Whitney, W. and Fergus, R · 2018
Later among the works it cites.
Task-agnostic dynamics priors for deep reinforcement learning
Du, Y. and Narasimhan, K · 2019
Closest in time.
Model-based reinforcement learning for Atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowsi, P., Levine, S., Sepassi, R., Tucker, G., and Michalewski, H · 2019
Closest in time.
Algorithmic framework for model-based deep reinforcement learning with theoretical guarantees
Luo, Y., Xu, H., Li, Y., Tian, Y., Darrell, T., and Ma, T · 2019
Closest in time.
Probabilistic planning with sequential Monte Carlo methods
Piché, A., Thomas, V., Ibrahim, C., Bengio, Y., and Pal, C · 2019
Closest in time.