Fetching the paper…
Reading the bibliography…
When environmental interaction is expensive, model-based reinforcement learning offers a solution by planning ahead and avoiding costly mistakes.
Science and statistics
Box, G. E. (1976) · 1976
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
Barto, A. G., Sutton, R. S., and Anderson, C. W. (1983) · 1983
Earlier work this paper cites.
Integrated modeling and control based on reinforcement learning
Sutton, R. S. (1990) · 1990
Earlier work this paper cites.
Td models: Modeling the world at a mixture of time scales
Sutton, R. S. (1995) · 1995
Earlier work this paper cites.
Reinforcement learning - an introduction
Sutton, R. S. and Barto, A. G. (1998) · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S. (1999) · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
Uncertainty-based competition between prefrontal and dorsolateral striatal systems for behavioral control
Daw, N. D., Niv, Y., and Dayan, P. (2005) · 2005
Earlier work this paper cites.
Using inaccurate models in reinforcement learning
Abbeel, P., Quigley, M., and Ng, A. Y. (2006) · 2006
Earlier work this paper cites.
Bandit based monte-carlo planning
Kocsis, L. and Szepesvári, C. (2006) · 2006
Earlier work this paper cites.
An analysis of linear models, linear value-function approximation, and feature selection for reinforcement learning
Parr, R., Li, L., Taylor, G., Painter-Wakefield, C., and Littman, M. L. (2008) · 2008
Earlier work this paper cites.
Dyna-style planning with linear function approximation and prioritized sweeping
Sutton, R. S., Szepesvári, C., Geramifard, A., and Bowling, M. H. (2008) · 2008
Earlier work this paper cites.
Multi-step dyna planning for policy evaluation and control
Yao, H., Bhatnagar, S., Diao, D., Sutton, R. S., and Szepesvári, C. (2009) · 2009
Earlier work this paper cites.
Linear options
Sorg, J. and Singh, S. (2010) · 2010
Cited alongside, same era.
PILCO: A model-based and data-efficient approach to policy search
Deisenroth, M. P. and Rasmussen, C. E. (2011) · 2011
Cited alongside, same era.
Retrospective revaluation in sequential decision making: A tale of two systems
Gershman, S. J., Markman, A. B., and Otto, A. R. (2014) · 2014
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L. (2014) · 2014
Cited alongside, same era.
Model regularization for stable sample rollouts
Talvitie, E. (2014) · 2014
Cited alongside, same era.
Strengths, weaknesses, and combinations of model-based and model-free reinforcement learning
Asadi, K. (2015) · 2015
Cited alongside, same era.
Improving pilco with bayesian neural network dynamics models
Gal, Y., McAllister, R., and Rasmussen, C. E. (2016) · 2016
Later among the works it cites.
Policy error bounds for model-based reinforcement learning with factored linear models
Pires, B. Á. and Szepesvári, C. (2016) · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T. P., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D. (2016) · 2016
Later among the works it cites.
Asadi, K., Allen, C., Roderick, M., Mohamed, A.-r., Konidaris, G., and Littman, M. (2017) · 2017
Later among the works it cites.
Value-aware loss function for model-based reinforcement learning
Farahmand, A.-m., Barreto, A., and Nikovski, D. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The dependence of effective planning horizon on model accuracy
Jiang, N., Kulesza, A., Singh, S., and Lewis, R. (2015) · 2015
Cited alongside, same era.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G. (2015) · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M. A., Fidjeland, A., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D. (2015) · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
Oh, J., Guo, X., Lee, H., Lewis, R. L., and Singh, S. (2015) · 2015
Cited alongside, same era.
A deeper look at planning as learning from replay
Vanseijen, H. and Sutton, R. (2015) · 2015
Cited alongside, same era.
Improving multi-step prediction of learned time series models
Venkatraman, A., Hebert, M., and Bagnell, J. A. (2015) · 2015
Cited alongside, same era.
Value prediction network
Oh, J., Singh, S., and Lee, H. (2017) · 2017
Later among the works it cites.
The predictron: End-to-end learning and planning
Silver, D., van Hasselt, H., Hessel, M., Schaul, T., Guez, A., Harley, T., Dulac-Arnold, G., Reichert, D. P., Rabinowitz, N. C., Barreto, A., and Degris, T. (2017) · 2017
Later among the works it cites.
Self-correcting models for model-based reinforcement learning
Talvitie, E. (2017) · 2017
Later among the works it cites.
Lipschitz continuity in model-based reinforcement learning
Asadi, K., Misra, D., and Littman, M. L. (2018b) · 2018
Closest in time.
Model-based value estimation for efficient model-free reinforcement learning
Feinberg, V., Wan, A., Stoica, I., Jordan, M. I., Gonzalez, J. E., and Levine, S. (2018) · 2018
Closest in time.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M. (2018) · 2018
Closest in time.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Nagabandi, A., Kahn, G., Fearing, R. S., and Levine, S. (2018) · 2018
Closest in time.