Fetching the paper…
Reading the bibliography…
Model-based Reinforcement Learning (MBRL) allows data-efficient learning which is required in real world applications such as robotics.
Model predictive control: theory and practice—a survey
Garcia, C. E., Prett, D. M., and Morari, M · 1989
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, R. S · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Introduction to reinforcement learning , volume 135
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
The cross-entropy method for combinatorial and continuous optimization
Rubinstein, R · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
Model selection and multimodel inference: a practical information-theoretic approach
Burnham, K. P. and Anderson, D. R · 2003
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Bertsekas, D. P., Bertsekas, D. P., Bertsekas, D. P., and Bertsekas, D. P · 2005
Earlier work this paper cites.
Robust constrained model predictive control
Richards, A. G · 2005
Earlier work this paper cites.
A generalized iterative lqg method for locally-optimal feedback control of constrained nonlinear stochastic systems
Todorov, E. and Li, W · 2005
Earlier work this paper cites.
Sample-based learning and search with permanent and transient memories
Silver, D., Sutton, R. S., and Müller, M · 2008
Earlier work this paper cites.
Dyna-style planning with linear function approximation and prioritized sweeping
Sutton, R. S., Szepesvári, C., Geramifard, A., and Bowling, M · 2008
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Deisenroth, M. and Rasmussen, C. E · 2011
Earlier work this paper cites.
Miscellaneous clustering methods
Everitt, B. S., Landau, S., Leese, M., and Stahl, D · 2011
Earlier work this paper cites.
Synthesis and stabilization of complex behaviors through online trajectory optimization
Tassa, Y., Erez, T., and Todorov, E · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Cited alongside, same era.
Guided policy search
Levine, S. and Koltun, V · 2013
Cited alongside, same era.
Learning neural network policies with guided policy search under unknown dynamics
Levine, S. and Abbeel, P · 2014
Cited alongside, same era.
Control-limited differential dynamic programming
Tassa, Y., Mansard, N., and Todorov, E · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V. et al · 2015
Cited alongside, same era.
Trust region policy optimization
Path integral guided policy search
Chebotar, Y., Kalakrishnan, M., Yahya, A., Li, A., Schaal, S., and Levine, S · 2017
Later among the works it cites.
Uncertainty-driven imagination for continuous deep reinforcement learning
Kalweit, G. and Boedecker, J · 2017
Later among the works it cites.
Value prediction network
Oh, J., Singh, S., and Lee, H · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Racanière, S., Weber, T., Reichert, D., Buesing, L., Guez, A., Jimenez Rezende, D., Puigdomènech Badia, A., Vinyals, O., Heess, N., Li, Y., Pascanu, R., Battaglia, P., Hassabis, D., Silver, D., and Wierstra, D · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
Information theoretic mpc for model-based reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
Openai gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Continuous deep q-learning with model-based acceleration
Gu, S., Lillicrap, T., Sutskever, I., and Levine, S · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V. et al · 2016
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P · 2016
Cited alongside, same era.
Williams, G., Wagener, N., Goldfain, B., Drews, P., Rehg, J. M., Boots, B., and Theodorou, E. A · 2017
Later among the works it cites.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Buckman, J., Hafner, D., Tucker, G., Brevdo, E., and Lee, H · 2018
Later among the works it cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S · 2018
Later among the works it cites.
Model-based reinforcement learning via meta-policy optimization
Clavera, I., Rothfuss, J., Schulman, J., Fujita, Y., Asfour, T., and Abbeel, P · 2018
Later among the works it cites.
Model-based value estimation for efficient model-free reinforcement learning
Feinberg, V., Wan, A., Stoica, I., Jordan, M. I., Gonzalez, J. E., and Levine, S · 2018
Later among the works it cites.
Data-efficient reinforcement learning with probabilistic model predictive control
Kamthe, S. and Deisenroth, M · 2018
Later among the works it cites.
Model-ensemble trust-region policy optimization
Kurutach, T., Clavera, I., Duan, Y., Tamar, A., and Abbeel, P · 2018
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Nagabandi, A., Kahn, G., Fearing, R. S., and Levine, S · 2018
Later among the works it cites.
Temporal difference models: Model-free deep RL for model-based control
Pong*, V., Gu*, S., Dalal, M., and Levine, S · 2018
Later among the works it cites.
Plan online, learn offline: Efficient learning and exploration via model-based control
Lowrey, K., Rajeswaran, A., Kakade, S., Todorov, E., and Mordatch, I · 2019
Closest in time.