Fetching the paper…
Reading the bibliography…
Model-free deep reinforcement learning algorithms have been shown to be capable of learning a wide range of robotic skills, but typically require a very large number of samples to achieve good performance.
R. Sutton, “ Dyna, an integrated architecture for learning, planning, and reacting ,” in AAAI , 1991
1991
Earlier work this paper cites.
K. J. Hunt, D. Sbarbaro, R. Żbikowski, and P. J. Gawthrop, “ Neural networks for control systems—a survey ,” in Automatica , 1992
1992
Earlier work this paper cites.
G. Bekey and K. Y. Goldberg, Neural networks in robotics . Springer US, 1992
1992
Earlier work this paper cites.
A. Y. Ng, S. J. Russell, et al. , “ Algorithms for inverse reinforcement learning ,” in ICML , 2000
2000
Earlier work this paper cites.
J. Morimoto and C. G. Atkeson, “Minimax differential dynamic programming: An application to robust biped walking,” in NIPS , 2003
2003
Earlier work this paper cites.
A. Richards, “ Robust constrained model predictive control ,” Ph.D. dissertation, MIT, 2004
2004
Earlier work this paper cites.
W. Li and E. Todorov, “ Iterative linear quadratic regulator design for nonlinear biological movement systems ,” in ICINCO , 2004
2004
Earlier work this paper cites.
J. Ko and D. Fox, “ GP-BayesFilters: Bayesian filtering using gaussian process prediction and observation models ,” in IROS , 2008
2008
Earlier work this paper cites.
D. Silver, R. S. Sutton, and M. Müller, “ Sample-based learning and search with permanent and transient memories ,” in ICML , 2008
2008
Earlier work this paper cites.
A. Rao, “ A survey of numerical methods for optimal control ,” in Advances in the Astronautical Sciences , 2009
2009
Earlier work this paper cites.
M. Deisenroth and C. Rasmussen, “ A model-based and data-efficient approach to policy search ,” in ICML , 2011
2011
Earlier work this paper cites.
M. P. Deisenroth, C. E. Rasmussen, and D. Fox, “Learning to control a low-cost manipulator using data-efficient reinforcement learning,” 2011
2011
Earlier work this paper cites.
S. M. Khansari-Zadeh and A. Billard, “ Learning stable nonlinear dynamical systems with gaussian mixture models ,” in IEEE Transactions on Robotics , 2011
2011
Earlier work this paper cites.
S. Ross, G. J. Gordon, and D. Bagnell, “ A reduction of imitation learning and structured prediction to no-regret online learning ,” in AISTATS , 2011
2011
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “ Mujoco: A physics engine for model-based control ,” in IROS , 2012
2012
Earlier work this paper cites.
M. P. Deisenroth, R. Calandra, A. Seyfarth, and J. Peters, “Toward fast policy search for learning legged locomotion,” in IROS , 2012
2012
Cited alongside, same era.
I. Grondman, L. Busoniu, G. A. Lopes, and R. Babuska, “ A survey of actor-critic reinforcement learning: standard and natural policy gradients ,” in IEEE Transactions on Systems, Man, and Cybernetics , 2012
2012
Cited alongside, same era.
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller, “ Playing Atari with deep reinforcement learning ,” in Workshop on Deep Learning, NIPS , 2013
2013
Cited alongside, same era.
M. P. Deisenroth, G. Neumann, J. Peters, et al. , “ A survey on policy search for robotics ,” in Foundations and Trends in Robotics , 2013
2013
Cited alongside, same era.
J. Kober, J. A. Bagnell, and J. Peters, “ Reinforcement learning in robotics: A survey ,” IJRR , 2013
2015
Later among the works it cites.
K. Asadi, “ Strengths, weaknesses, and combinations of model-based and model-free reinforcement learning ,” Ph.D. dissertation, University of Alberta, 2015
2015
Later among the works it cites.
N. Heess, G. Wayne, D. Silver, T. Lillicrap, T. Erez, and Y. Tassa, “Learning continuous control policies by stochastic value gradients,” in NIPS , 2015
2015
Later among the works it cites.
J. Oh, V. Chockalingam, S. Singh, and H. Lee, “ Control of memory, active perception, and action in minecraft ,” in ICML , 2016
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2013
Cited alongside, same era.
R. Lioutikov, A. Paraschos, J. Peters, and G. Neumann, “ Sample-based information-theoretic stochastic optimal control ,” in ICRA , 2014
2014
Cited alongside, same era.
J. Boedecker, J. T. Springenberg, J. Wülfing, and M. Riedmiller, “ Approximate real-time optimal control based on sparse gaussian process models ,” in ADPRL , 2014
2014
Cited alongside, same era.
S. Levine and P. Abbeel, “ Learning neural network policies with guided policy search under unknown dynamics ,” in NIPS , 2014
2014
Cited alongside, same era.
M. C. Yip and D. B. Camarillo, “ Model-less feedback control of continuum manipulators in constrained environments ,” in IEEE Transactions on Robotics , 2014
2014
Cited alongside, same era.
D. Kingma and J. Ba, “ Adam: A method for stochastic optimization ,” in ICLR , 2014
2014
Cited alongside, same era.
J. Schulman, S. Levine, P. Moritz, M. Jordan, and P. Abbeel, “ Trust region policy optimization ,” in ICML , 2015
2015
Cited alongside, same era.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al. , “ Human-level control through deep reinforcement learning ,” in Nature , 2015
2015
Cited alongside, same era.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
Y. Gal, R. T. McAllister, and C. E. Rasmussen, “ Improving PILCO with bayesian neural network dynamics models ,” in Data-Efficient Machine Learning workshop , 2016
2016
Later among the works it cites.
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel, “ Benchmarking deep reinforcement learning for continuous control ,” in ICML , 2016
2016
Later among the works it cites.
N. Mishra, P. Abbeel, and I. Mordatch, “Prediction and control with temporal segment models,” in ICML , 2017
2017
Closest in time.
2017
Closest in time.
S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,” in JMLR , 2017
2017
Closest in time.