Fetching the paper…
Reading the bibliography…
Deep learning and reinforcement learning methods have recently been used to solve a variety of problems in continuous control domains.
Neural networks for control
Paul J. Webros · 1990
Earlier work this paper cites.
Neural networks for control systems: A survey
K. J. Hunt, D. Sbarbaro, R. Żbikowski, and P. J. Gawthrop · 1992
Earlier work this paper cites.
A direct adaptive method for faster backpropagation learning: The RPROP algorithm
M. Riedmiller and H. Braun · 1993
Earlier work this paper cites.
Temporal difference learning and td-gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Autonomous helicopter control using reinforcement learning policy search methods
J Andrew Bagnell and Jeff G Schneider · 2001
Earlier work this paper cites.
Policy gradient reinforcement learning for fast quadrupedal locomotion
Nate Kohl and Peter Stone · 2004
Earlier work this paper cites.
Neural fitted Q iteration - first experiences with a data efficient neural reinforcement learning method
Martin A. Riedmiller · 2005
Earlier work this paper cites.
Learning cpg-based biped locomotion with a policy gradient method
Takamitsu Matsubara, Jun Morimoto, Jun Nakanishi, Masa-aki Sato, and Kenji Doya · 2006
Earlier work this paper cites.
Policy gradient methods for robotics
Jan Peters and Stefan Schaal · 2006
Earlier work this paper cites.
Dynamic Movement Primitives -A Framework for Motor Control in Humans and Humanoid Robotics , pages 261–280
Stefan Schaal · 2006
Earlier work this paper cites.
Neural reinforcement learning controllers for a real robot application
Roland Hafner and Martin A. Riedmiller · 2007
Earlier work this paper cites.
Relative entropy inverse reinforcement learning
A. Boularias, J. Kober, and J. Peters · 2011
Cited alongside, same era.
Reinforcement learning in feedback control
Roland Hafner and Martin Riedmiller · 2011
Cited alongside, same era.
Learning force control policies for compliant manipulation
M. Kalakrishnan, L. Righetti, P. Pastor, and S. Schaal · 2011
Cited alongside, same era.
Skill learning and task outcome prediction for manipulation
P. Pastor, M. Kalakrishnan, S. Chitta, E. Theodorou, and S. Schaal · 2011
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
A survey on policy search for robotics
Marc Peter Deisenroth, Gerhard Neumann, Jan Peters, et al · 2013
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Later among the works it cites.
Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours
Lerrel Pinto and Abhinav Gupta · 2015
Later among the works it cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael I. Jordan, and Philipp Moritz · 2015
Later among the works it cites.
Learning robot in-hand manipulation with tactile features
Herke van Hoof, Tucker Hermans, Gerhard Neumann, and Jan Peters · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning objective functions for manipulation
M. Kalakrishnan, P. Pastor, L. Righetti, and S. Schaal · 2013
Cited alongside, same era.
Learning to select and generalize striking movements in robot table tennis
K. Muelling, J. Kober, O. Kroemer, and J. Peters · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Learning neural network policies with guided policy search under unknown dynamics
Sergey Levine and Pieter Abbeel · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Tim Lillicrap, Tom Erez, and Yuval Tassa · 2015
Cited alongside, same era.
Guided cost learning: Deep inverse optimal control via policy optimization
Chelsea Finn, Sergey Levine, and Pieter Abbeel · 2016
Later among the works it cites.
Learning dexterous manipulation for a soft robotic hand from human demonstrations
Abhishek Gupta, Clemens Eppner, Sergey Levine, and Pieter Abbeel · 2016
Later among the works it cites.
Sergey Levine, Peter Pastor, Alex Krizhevsky, and Deirdre Quillen · 2016
Later among the works it cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy P Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Later among the works it cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2016
Later among the works it cites.
Collective robot reinforcement learning with distributed asynchronous guided policy search
Ali Yahya, Adrian Li, Mrinal Kalakrishnan, Yevgen Chebotar, and Sergey Levine · 2016
Later among the works it cites.