Fetching the paper…
Reading the bibliography…
Model usage is the central challenge of model-based reinforcement learning.
Model-based reinforcement learning for atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., Sepassi, R., Tucker, G., and Michalewski, H · 1903
Earlier work this paper cites.
When to trust your model: Model-based policy optimization
Janner, M., Fu, J., Zhang, M., and Levine, S · 1906
Earlier work this paper cites.
Exploring model-based planning with policy networks
Wang, T. and Ba, J · 1906
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Sutton, R. S · 1991
Earlier work this paper cites.
Q-learning
Watkins, C. J. C. H. and Dayan, P · 1992
Earlier work this paper cites.
A natural policy gradient
Kakade, S · 2001
Earlier work this paper cites.
A tutorial on the cross-entropy method
de Boer, P.-T., Kroese, D. P., Mannor, S., and Rubinstein, R. Y · 2004
Earlier work this paper cites.
A survey of numerical methods for optimal control
Rao, A · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Deisenroth, M. P. and Rasmussen, C. E · 2011
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Cited alongside, same era.
Guided policy search
Levine, S. and Koltun, V · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. A · 2013
Cited alongside, same era.
Understanding Machine Learning: From Theory to Algorithms
Shalev-Shwartz, S. and Ben-David, S · 2014
Cited alongside, same era.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Nagabandi, A., Kahn, G., Fearing, R. S., and Levine, S · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Buckman, J., Hafner, D., Tucker, G., Brevdo, E., and Lee, H · 2018
Later among the works it cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S · 2018
Later among the works it cites.
Model-based value estimation for efficient model-free reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Schulman, J., Levine, S., Moritz, P., Jordan, M. I., and Abbeel, P · 2015
Cited alongside, same era.
Model predictive path integral control using covariance variable importance sampling
Williams, G., Aldrich, A., and Theodorou, E. A · 2015
Cited alongside, same era.
Prioritized experience replay, 2015
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2016
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S
Cited in the paper.
Soft actor-critic algorithms and applications
Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., and Levine, S
Cited in the paper.
Feinberg, V., Wan, A., Stoica, I., Jordan, M. I., Gonzalez, J. E., and Levine, S · 2018
Later among the works it cites.
Model-ensemble trust-region policy optimization
Kurutach, T., Clavera, I., Duan, Y., Tamar, A., and Abbeel, P · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
Algorithmic framework for model-based reinforcement learning with theoretical guarantees
Xu, H., Li, Y., Tian, Y., Darrell, T., and Ma, T · 2018
Later among the works it cites.