Fetching the paper…
Reading the bibliography…
A key component of model-based reinforcement learning (RL) is a dynamics model that predicts the outcomes of actions.
Model predictive heuristic control: Applications to industial processes
Testud, J., Richalet, J., Rault, A., and Papon, J · 1978
Earlier work this paper cites.
Neural network modeling and an extended dmc algorithm to control nonlinear systems
Hernandaz, E. P. S. and Arkun, Y · 1990
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Sutton, R. S · 1991
Earlier work this paper cites.
Model predictive control using neural networks
Draeger, A., Engell, S., and Ranke, H. D · 1995
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Learning options in reinforcement learning
Stolle, M. and Precup, D · 2002
Earlier work this paper cites.
The IM algorithm: a variational approach to information maximization
Barber, D. and Agakov, F · 2003
Earlier work this paper cites.
Gaussian processes in reinforcement learning
Rasmussen, C. E. and Kuss, M · 2003
Earlier work this paper cites.
Towards a unified theory of state abstraction for mdps
Li, L., Walsh, T. J., and Littman, M. L · 2006
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Deisenroth, M. P. and Rasmussen, C. E · 2011
Earlier work this paper cites.
Agnostic system identification for model-based reinforcement learning
Ross, S. and Bagnell, J. A · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
Heess, N. M. O., Wayne, G., Silver, D., Lillicrap, T. P., Erez, T., and Tassa, Y · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Deepmpc: Learning deep latent features for model predictive control
Lenz, I., Knepper, R. A., and Saxena, A · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Watter, M., Springenberg, J. T., Boedecker, J., and Riedmiller, M. A · 2015
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Gregor, K., Rezende, D. J., and Wierstra, D · 2016
Earlier work this paper cites.
Vime: Variational information maximizing exploration
Houthooft, R., Chen, X., Duan, Y., Schulman, J., Turck, F. D., and Abbeel, P · 2016
Earlier work this paper cites.
Aggressive driving with model predictive path integral control
Williams, G., Drews, P., Goldfain, B., Rehg, J. M., and Theodorou, E. A · 2016
Earlier work this paper cites.
The option-critic architecture
Bacon, P.-L., Harb, J., and Precup, D · 2017
Earlier work this paper cites.
A laplacian framework for option discovery in reinforcement learning
Machado, M. C., Bellemare, M. G., and Bowling, M · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Earlier work this paper cites.
Epopt: Learning robust neural network policies using model ensembles
Rajeswaran, A., Ghotra, S., Levine, S., and Ravindran, B · 2017
Earlier work this paper cites.
Feudal networks for hierarchical reinforcement learning
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N. M. O., Jaderberg, M., Silver, D., and Kavukcuoglu, K · 2017
Earlier work this paper cites.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Buckman, J., Hafner, D., Tucker, G., Brevdo, E., and Lee, H · 2018
Earlier work this paper cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S · 2018
Cited alongside, same era.
Model-based reinforcement learning via meta-policy optimization
Clavera, I., Rothfuss, J., Schulman, J., Fujita, Y., Asfour, T., and Abbeel, P · 2018
Cited alongside, same era.
Visual foresight: Model-based deep reinforcement learning for vision-based robotic control
Ebert, F., Finn, C., Dasari, S., Xie, A., Lee, A. X., and Levine, S · 2018
Cited alongside, same era.
Model-based value expansion for efficient model-free reinforcement learning
Feinberg, V., Wan, A., Stoica, I., Jordan, M. I., Gonzalez, J. E., and Levine, S · 2018
Cited alongside, same era.
Recurrent world models facilitate policy evolution
Ha, D. and Schmidhuber, J · 2018
Cited alongside, same era.
Mastering Atari, Go, Chess and Shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., Lillicrap, T., and Silver, D · 2020
Later among the works it cites.
Planning to explore via self-supervised world models
Sekar, R., Rybkin, O., Daniilidis, K., Abbeel, P., Hafner, D., and Pathak, D · 2020
Later among the works it cites.
Dynamics-aware unsupervised discovery of skills
Sharma, A., Gu, S., Levine, S., Kumar, V., and Hausman, K · 2020
Later among the works it cites.
Latent skill planning for exploration and transfer
Xie, K., Bharadhwaj, H., Hafner, D., Garg, A., and Shkurti, F · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J. Y., Levine, S., Finn, C., and Ma, T · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kurutach, T., Clavera, I., Duan, Y., Tamar, A., and Abbeel, P · 2018
Cited alongside, same era.
Data-efficient hierarchical reinforcement learning
Nachum, O., Gu, S. S., Lee, H., and Levine, S · 2018
Cited alongside, same era.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A. J., and Klimov, O · 2019
Cited alongside, same era.
Scalable methods for computing state similarity in deterministic markov decision processes
Castro, P. S · 2019
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2019
Cited alongside, same era.
Deepmdp: Learning continuous latent space models for representation learning
Gelada, C., Kumar, S., Buckman, J., Nachum, O., and Bellemare, M. G · 2019
Cited alongside, same era.
Learning latent dynamics for planning from pixels
Hafner, D., Lillicrap, T. P., Fischer, I. S., Villegas, R., Ha, D. R., Lee, H., and Davidson, J · 2019
Cited alongside, same era.
Opal: Offline primitive discovery for accelerating offline reinforcement learning
Ajay, A., Kumar, A., Agrawal, P., Levine, S., and Nachum, O · 2021
Later among the works it cites.
Model-based offline planning
Argenson, A. and Dulac-Arnold, G · 2021
Later among the works it cites.
Smirl: Surprise minimizing reinforcement learning in unstable environments
Berseth, G., Geng, D., Devin, C., Rhinehart, N., Finn, C., Jayaraman, D., and Levine, S · 2021
Later among the works it cites.
Robust predictable control
Eysenbach, B., Salakhutdinov, R., and Levine, S · 2021
Later among the works it cites.
Unsupervised skill discovery with bottleneck option learning
Kim, J., Park, S., and Kim, G · 2021
Later among the works it cites.
Behavior from the void: Unsupervised active pre-training
Liu, H. and Abbeel, P · 2021
Later among the works it cites.
Reset-free lifelong learning with skill-space planning
Lu, K., Grover, A., Abbeel, P., and Mordatch, I · 2021
Later among the works it cites.
Discovering and achieving goals via world models
Mendonca, R., Rybkin, O., Daniilidis, K., Hafner, D., and Pathak, D · 2021
Later among the works it cites.
Temporal predictive coding for model-based planning in latent space
Nguyen, T. D., Shu, R., Pham, T., Bui, H. H., and Ermon, S · 2021
Later among the works it cites.
Information is power: Intrinsic control via information capture
Rhinehart, N., Wang, J., Berseth, G., Co-Reyes, J. D., Hafner, D., Finn, C., and Levine, S · 2021
Later among the works it cites.
Data-efficient hindsight off-policy option learning
Wulfmeier, M., Rao, D., Hafner, R., Lampe, T., Abdolmaleki, A., Hertweck, T., Neunert, M., Tirumala, D., Siegel, N., Heess, N. M. O., and Riedmiller, M. A · 2021
Later among the works it cites.
Reinforcement learning with prototypical representations
Yarats, D., Fergus, R., Lazaric, A., and Pinto, L · 2021
Later among the works it cites.
Temporal difference learning for model predictive control
Hansen, N., Wang, X., and Su, H · 2022
Later among the works it cites.
Lyapunov density models: Constraining distribution shift in learning-based control
Kang, K., Gradu, P., Choi, J. J., Janner, M., Tomlin, C. J., and Levine, S · 2022
Later among the works it cites.
Unsupervised reinforcement learning with contrastive intrinsic control
Laskin, M., Liu, H., Peng, X. B., Yarats, D., Rajeswaran, A., and Abbeel, P · 2022
Later among the works it cites.
Lipschitz-constrained unsupervised skill discovery
Park, S., Choi, J., Kim, J., Lee, H., and Kim, G · 2022
Later among the works it cites.
Unsupervised model-based pre-training for data-efficient control from pixels
Rajeswar, S., Mazzaglia, P., Verbelen, T., Pich’e, A., Dhoedt, B., Courville, A. C., and Lacoste, A · 2022
Later among the works it cites.
Mo2: Model-based offline options
Salter, S., Wulfmeier, M., Tirumala, D., Heess, N. M. O., Riedmiller, M. A., Hadsell, R., and Rao, D · 2022
Later among the works it cites.
Learning off-policy with online planning
Sikchi, H. S., Zhou, W., and Held, D · 2022
Later among the works it cites.
Learning more skills through optimistic exploration
Strouse, D., Baumli, K., Warde-Farley, D., Mnih, V., and Hansen, S. S · 2022
Later among the works it cites.
Daydreamer: World models for physical robot learning
Wu, P., Escontrela, A., Hafner, D., Goldberg, K., and Abbeel, P · 2022
Later among the works it cites.