Fetching the paper…
Reading the bibliography…
Data-driven model predictive control has two key advantages over model-free methods: a potential for improved sample efficiency through model learning, and better performance as computational budget for planning increases.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T. P., Ba, J., and Norouzi, M · 1912
Earlier work this paper cites.
Learning to predict by the method of temporal differences
Sutton, R · 1988
Earlier work this paper cites.
Optimization of computer simulation models with rare events
Rubinstein, R. Y · 1997
Earlier work this paper cites.
Learning-based model predictive control for markov decision processes
Negenborn, R. R., De Schutter, B., Wiering, M. A., and Hellendoorn, H · 2005
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R · 2005
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A., Bagnell, J. A., and Dey, A. K · 2008
Earlier work this paper cites.
Mastering atari with discrete world models
Hafner, D., Lillicrap, T., Norouzi, M., and Ba, J · 2010
Earlier work this paper cites.
Synthesis and stabilization of complex behaviors through online trajectory optimization
Tassa, Y., Erez, T., and Todorov, E · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Model predictive path integral control using covariance variable importance sampling
Williams, G., Aldrich, A., and Theodorou, E. A · 2015
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Hasselt, H. V., Guez, A., and Silver, D · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T., Hunt, J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2016
Earlier work this paper cites.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Buckman, J., Hafner, D., Tucker, G., Brevdo, E., and Lee, H · 2018
Earlier work this paper cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S · 2018
Earlier work this paper cites.
Visual foresight: Model-based deep reinforcement learning for vision-based robotic control
Ebert, F., Finn, C., Dasari, S., Xie, A., Lee, A. X., and Levine, S · 2018
Earlier work this paper cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., Legg, S., and Kavukcuoglu, K · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H. V., and Meger, D · 2018
Earlier work this paper cites.
Recurrent world models facilitate policy evolution
Ha, D. and Schmidhuber, J · 2018
Earlier work this paper cites.
Soft actor-critic algorithms and applications
Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., and Levine, S · 2018
Cited alongside, same era.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Nagabandi, A., Kahn, G., Fearing, R. S., and Levine, S · 2018
Cited alongside, same era.
Temporal difference models: Model-free deep rl for model-based control
Pong, V. H., Gu, S. S., Dalal, M., and Levine, S · 2018
Cited alongside, same era.
Deepmind control suite
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., de Las Casas, D., Budden, D., Abdolmaleki, A., et al · 2018
Grac: Self-guided and self-regularized actor-critic
Shao, L., You, Y., Yan, M., Sun, Q., and Bohg, J · 2020
Later among the works it cites.
Curl: Contrastive unsupervised representations for reinforcement learning
Srinivas, A., Laskin, M., and Abbeel, P · 2020
Later among the works it cites.
Exploring model-based planning with policy networks
Wang, T. and Ba, J · 2020
Later among the works it cites.
Soft actor-critic (sac) implementation in pytorch
Yarats, D. and Kostrikov, I · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J. Y., Levine, S., Finn, C., and Ma, T · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Solar: Deep structured latent representations for model-based reinforcement learning
Zhang, M., Vikram, S., Smith, L., Abbeel, P., Johnson, M. J., and Levine, S · 2018
Cited alongside, same era.
Curious meta-controller: Adaptive alternation between model-based and model-free control in deep reinforcement learning
Hafez, M. B., Weber, C., Kerzel, M., and Wermter, S · 2019
Cited alongside, same era.
Learning latent dynamics for planning from pixels
Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., and Davidson, J · 2019
Cited alongside, same era.
When to trust your model: Model-based policy optimization
Janner, M., Fu, J., Zhang, M., and Levine, S · 2019
Cited alongside, same era.
Plan online, learn offline: Efficient learning and exploration via model-based control
Lowrey, K., Rajeswaran, A., Kakade, S. M., Todorov, E., and Mordatch, I · 2019
Cited alongside, same era.
Cem-rl: Combining evolutionary and gradient-based methods for policy search
Pourchot, A. and Sigaud, O · 2019
Cited alongside, same era.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2019
Cited alongside, same era.
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A., and Bellemare, M. G · 2021
Later among the works it cites.
Argenson, A. and Dulac-Arnold, G · 2021
Later among the works it cites.
Blending mpc & value function approximation for efficient reinforcement learning
Bhardwaj, M., Choudhury, S., and Boots, B · 2021
Later among the works it cites.
Exploring simple siamese representation learning
Chen, X. and He, K · 2021
Later among the works it cites.
Generalization in reinforcement learning by soft data augmentation
Hansen, N. and Wang, X · 2021
Later among the works it cites.
Self-supervised policy adaptation during deployment
Hansen, N., Jangir, R., Sun, Y., Alenyà, G., Abbeel, P., Efros, A. A., Pinto, L., and Wang, X · 2021
Later among the works it cites.
The value of planning for infinite-horizon model predictive control
Hatch, N. and Boots, B · 2021
Later among the works it cites.
Mt-opt: Continuous multi-task robotic reinforcement learning at scale
Kalashnikov, D., Varley, J., Chebotar, Y., Swanson, B., Jonschkowski, R., Finn, C., Levine, S., and Hausman, K · 2021
Later among the works it cites.
Learning to jump from pixels
Margolis, G., Chen, T., Paigwar, K., Fu, X., Kim, D., Kim, S., and Agrawal, P · 2021
Later among the works it cites.
Model predictive actor-critic: Accelerating robot skill acquisition with deep reinforcement learning
Morgan, A. S., Nandha, D., Chalvatzaki, G., D’Eramo, C., Dollar, A. M., and Peters, J · 2021
Later among the works it cites.
Temporal predictive coding for model-based planning in latent space
Nguyen, T. D., Shu, R., Pham, T., Bui, H. H., and Ermon, S · 2021
Later among the works it cites.
Mastering visual continuous control: Improved data-augmented reinforcement learning
Yarats, D., Fergus, R., Lazaric, A., and Pinto, L · 2021
Later among the works it cites.
Mastering atari games with limited data
Ye, W., Liu, S., Kurutach, T., Abbeel, P., and Gao, Y · 2021
Later among the works it cites.
Learning off-policy with online planning
Sikchi, H., Zhou, W., and Held, D · 2022
Closest in time.