Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) algorithms for real-world robotic applications need a data-efficient learning process and the ability to handle complex, unknown dynamical systems.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, R · 1990
Earlier work this paper cites.
Control, planning, learning, and imitation with dynamic movement primitives
Schaal, S., Peters, J., Nakanishi, J., and Ijspeert, A · 2003
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
Peters, J. and Schaal, S · 2008
Earlier work this paper cites.
Learning and generalization of motor skills by learning from demonstration
Pastor, P., Hoffmann, H., Asfour, T., and Schaal, S · 2009
Earlier work this paper cites.
Relative entropy policy search
Peters, J., Mülling, K., and Altun, Y · 2010
Earlier work this paper cites.
A generalized path integral control approach to reinforcement learning
Theodorou, E., Buchli, J., and Schaal, S · 2010
Earlier work this paper cites.
Learning to control a low-cost manipulator using data-efficient reinforcement learning
Deisenroth, M., Rasmussen, C., and Fox, D · 2011
Earlier work this paper cites.
Synthesis and stabilization of complex behaviors
Tassa, Y., Erez, T., and Todorov, E · 2012
Earlier work this paper cites.
Learning sequential motor tasks
Daniel, Christian, Neumann, Gerhard, Kroemer, Oliver, and Peters, Jan · 2013
Earlier work this paper cites.
A survey on policy search for robotics
Deisenroth, M., Neumann, G., and Peters, J · 2013
Earlier work this paper cites.
Auto-encoding variational Bayes
Kingma, D. and Welling, M · 2013
Cited alongside, same era.
Reinforcement learning in robotics: a survey
Kober, J., Bagnell, J., and Peters, J · 2013
Cited alongside, same era.
Playing Atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Cited alongside, same era.
Gaussian processes for data-efficient learning in robotics and control
Deisenroth, M., Fox, D., and Rasmussen, C · 2014
Cited alongside, same era.
Learning of closed-loop motion control
Farshidian, F., Neunert, M., and Buchli, J · 2014
Cited alongside, same era.
Learning neural network policies with guided policy search under unknown dynamics
Levine, S. and Abbeel, P · 2014
Trust region policy optimization
Schulman, J., Levine, S., Moritz, P., Jordan, M., and Abbeel, P · 2015
Later among the works it cites.
Model-free trajectory optimization for reinforcement learning
Akrour, R., Abdolmaleki, A., Abdulsamad, H., and Neumann, G · 2016
Later among the works it cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Later among the works it cites.
Continuous deep Q-learning with model-based acceleration
Gu, S., Lillicrap, T., Sutskever, I., and Levine, S · 2016
Later among the works it cites.
End-to-end training of deep visuomotor policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sample-based information-theoretic stochastic optimal control
Lioutikov, R., Paraschos, A., Neumann, G., and Peters, J · 2014
Cited alongside, same era.
Probabilistic differential dynamic programming
Pan, Y. and Theodorou, E · 2014
Cited alongside, same era.
Learning continuous control policies by stochastic value gradients
Heess, N., Wayne, G., Silver, D., Lillicrap, T., Tassa, Y., and Erez, T · 2015
Cited alongside, same era.
Learning contact-rich manipulation skills with guided policy search
Levine, S., Wagener, N., and Abbeel, P · 2015
Cited alongside, same era.
Lillicrap, T., Hunt, J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Later among the works it cites.
Guided policy search via approximate mirror descent
Montgomery, W. and Levine, S · 2016
Later among the works it cites.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P · 2016
Later among the works it cites.
Path integral guided policy search
Chebotar, Y., Kalakrishnan, M., Yahya, A., Li, A., Schaal, S., and Levine, S · 2017
Closest in time.
Reset-free guided policy search: efficient deep reinforcement learning with stochastic initial states
Montgomery, W., Ajay, A., Finn, C., Abbeel, P., and Levine, S · 2017
Closest in time.