Fetching the paper…
Reading the bibliography…
Model-based reinforcement learning (RL) has proven to be a data efficient approach for learning control tasks but is difficult to utilize in domains with complex observations such as images.
Differential Dynamic Programming
Jacobson, D. and Mayne, D · 1970
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, R · 1990
Earlier work this paper cites.
Reinforcement learning with soft state aggregation
Singh, S., Jaakkola, T., and Jordan, M · 1994
Earlier work this paper cites.
Applications of the self-organizing map to reinforcement learning
Smith, A · 2002
Earlier work this paper cites.
Convex Optimization
Boyd, S. and Vandenberghe, L · 2004
Earlier work this paper cites.
A generalized iterative LQG method for locally-optimal feedback control of constrained nonlinear stochastic systems
Todorov, E. and Li, W · 2005
Earlier work this paper cites.
Variational message passing
Winn, J. and Bishop, C · 2005
Earlier work this paper cites.
Deep auto-encoder neural networks in reinforcement learning
Lange, S. and Riedmiller, M · 2010
Earlier work this paper cites.
Dimension reduction and its application to model-based exploration in continuous spaces
Nouri, A. and Littman, M · 2010
Earlier work this paper cites.
Synthesis and stabilization of complex behaviors
Tassa, Y., Erez, T., and Todorov, E · 2012
Earlier work this paper cites.
Model Predictive Control
Camacho, E. and Alba, C · 2013
Earlier work this paper cites.
Stochastic variational inference
Hoffman, M., Blei, D., Wang, C., and Paisley, J · 2013
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Kober, J., Bagnell, J., and Peters, J · 2013
Earlier work this paper cites.
Gaussian processes for data-efficient learning in robotics and control
Deisenroth, M., Fox, D., and Rasmussen, C · 2014
Earlier work this paper cites.
Auto-encoding variational Bayes
Kingma, D. and Welling, M · 2014
Earlier work this paper cites.
Learning neural network policies with guided policy search under unknown dynamics
Levine, S. and Abbeel, P · 2014
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, D., Mohamed, S., and Wierstra, D · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2015
Cited alongside, same era.
Optimism-driven exploration for nonlinear systems
Moldovan, T., Levine, S., Jordan, M., and Abbeel, P · 2015
Cited alongside, same era.
Embed to control: A locally linear latent dynamics model for control from raw images
Watter, M., Springenberg, J., Boedecker, J., and Riedmiller, M · 2015
Cited alongside, same era.
Learning to poke by poking: Experiential learning of intuitive physics
Agrawal, P., Nair, A., Abbeel, P., Malik, J., and Levine, S · 2016
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
Deep reinforcement learning for tensegrity robot locomotion
Zhang, M., Geng, X., Bruce, J., Caluwaerts, K., Vespignani, M., SunSpiral, V., Abbeel, P., and Levine, S · 2017
Later among the works it cites.
Learning dexterous in-hand manipulation
Andrychowicz, M., Baker, B., Chociej, M., Józefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., Schneider, J., Sidor, S., Tobin, J., Welinder, P., Weng, L., and Zaremba, W · 2018
Closest in time.
Robust locally-linear controllable embedding
Banijamali, E., Shu, R., Ghavamzadeh, M., Bui, H., and Ghodsi, A · 2018
Closest in time.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Deep spatial autoencoders for visuomotor learning
Finn, C., Tan, X., Duan, Y., Darrell, T., Levine, S., and Abbeel, P · 2016
Cited alongside, same era.
Composing graphical models with neural networks for structured representations and fast inference
Johnson, M., Duvenaud, D., Wiltschko, A., Datta, S., and Adams, R · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P · 2016
Cited alongside, same era.
Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours
Pinto, L. and Gupta, A · 2016
Cited alongside, same era.
Combining model-based and model-free updates for trajectory-centric reinforcement learning
Chebotar, Y., Hausman, K., Zhang, M., Sukhatme, G., Schaal, S., and Levine, S · 2017
Cited alongside, same era.
Visual foresight: Model-based deep reinforcement learning for vision-based robotic control
Ebert, F., Finn, C., Dasari, S., Xie, A., Lee, A., and Levine, S · 2018
Closest in time.
Model-based value estimation for efficient model-free reinforcement learning
Feinberg, V., Wan, A., Stoica, I., Jordan, M., Gonzalez, J., and Levine, S · 2018
Closest in time.
Variational inverse control with events: A general framework for data-driven reward definition
Fu, J., Singh, A., Ghosh, D., Yang, L., and Levine, S · 2018
Closest in time.
Addressing function approximation error in actor-critic methods
Fujimoto, S., van Hoof, H., and Meger, D · 2018
Closest in time.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Closest in time.
Learning latent dynamics for planning from pixels
Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., and Davidson, J · 2018
Closest in time.
State representation learning for control: An overview
Lesort, T., Díaz-Rodríguez, N., Goudou, J., and Filliat, D · 2018
Closest in time.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Nagabandi, A., Kahn, G., Fearing, R., and Levine, S · 2018
Closest in time.
Dexterous manipulation with deep reinforcement learning: Efficient, general, and low-cost
Zhu, H., Gupta, A., Rajeswaran, A., Levine, S., and Kumar, V · 2018
Closest in time.