Fetching the paper…
Reading the bibliography…
In principle, reinforcement learning and policy search methods can enable robots to learn highly complex and general skills that may allow them to function amid the complexity and diversity of the real world.
A platform for robotics research based on the remote-brained robot approach
M. Inaba, S. Kagami, F. Kanehiro, and Y. Hoshino · 2000
Earlier work this paper cites.
Movement imitation with nonlinear dynamical systems in humanoid robots
J.A. Ijspeert, J. Nakanishi, and S. Schaal · 2002
Earlier work this paper cites.
Stochastic policy gradient reinforcement learning on a simple 3d biped
R. Tedrake, T.W. Zhang, and H.S. Seung · 2004
Earlier work this paper cites.
Iterative linear quadratic regulator design for nonlinear biological movement systems
W. Li and E. Todorov · 2004
Earlier work this paper cites.
Relative entropy policy search
J. Peters, K. Mülling, and Y. Altun · 2010
Earlier work this paper cites.
A generalized path integral control approach to reinforcement learning
E. Theodorou, J. Buchli, and S. Schaal · 2010
Earlier work this paper cites.
Cloud-enabled humanoid robots
J. Kuffner · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J. Smola · 2010
Earlier work this paper cites.
Hierarchical relative entropy policy search
Christian Daniel, Gerhard Neumann, and Jan Peters · 2012
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al · 2012
Cited alongside, same era.
A survey on policy search for robotics
M.P. Deisenroth, G. Neumann, and J. Peters · 2013
Cited alongside, same era.
Acquiring visual servoing reaching and grasping skills using neural reinforcement learning
Thomas Lampe and Martin Riedmiller · 2013
Cited alongside, same era.
Cloud-based robot grasping with the google object recognition engine
B. Kehoe, A. Matsukawa, S. Candido, J. Kuffner, and K. Goldberg · 2013
Cited alongside, same era.
Learning neural network policies with guided policy search under unknown dynamics
S. Levine and P. Abbeel · 2014
Cited alongside, same era.
Bregman alternating direction method of multipliers
H. Wang and A. Banerjee · 2014
A survey of research on cloud robotics and automation
B. Kehoe, S. Patil, P. Abbeel, and K. Goldberg · 2015
Later among the works it cites.
Interactive control of diverse complex characters with neural networks
Igor Mordatch, Kendall Lowrey, Galen Andrew, Zoran Popovic, and Emanuel V Todorov · 2015
Later among the works it cites.
TensorFlow: Large-scale machine learning on heterogeneous distributed systems, 2015
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2015
Later among the works it cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Closest in time.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Joint training of a convolutional network and a graphical model for human pose estimation
J. Tompson, A. Jain, Y. LeCun, and C. Bregler · 2014
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Moritz, M. Jordan, and P. Abbeel · 2015
Cited alongside, same era.
Deep learning
Yann Lecun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
Closest in time.
Guided policy search as approximate mirror descent
W. Montgomery and S. Levine · 2016
Closest in time.
Revisiting distributed synchronous sgd
Jianmin Chen, Rajat Monga, Samy Bengio, and Rafal Jozefowicz · 2016
Closest in time.
Going further with point pair features
S. Hinterstoisser, V. Lepetit, N. Rajkumar, and K. Konolige · 2016
Closest in time.
Path integral guided policy search
Y. Chebotar, M. Kalakrishnan, A. Yahya, A. Li, S. Schaal, and S. Levine · 2017
Closest in time.