Fetching the paper…
Reading the bibliography…
Policy search methods can allow robots to learn control policies for a wide range of tasks, but practical applications of policy search often require hand-engineered components for perception, state estimation, and low-level control.
Differential Dynamic Programming
D. Jacobson and D. Mayne · 1970
Earlier work this paper cites.
Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position
K. Fukushima · 1980
Earlier work this paper cites.
Handwritten digit recognition with a back-propagation network
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel · 1989
Earlier work this paper cites.
ALVINN: an autonomous land vehicle in a neural network
D. Pomerleau · 1989
Earlier work this paper cites.
A stochastic reinforcement learning algorithm for learning real-valued functions
V. Gullapalli · 1990
Earlier work this paper cites.
Neural Networks in Robotics
G. Bekey and K. Goldberg · 1992
Earlier work this paper cites.
A new approach to visual servoing in robotics
B. Espiau, F. Chaumette, and P. Rives · 1992
Earlier work this paper cites.
Neural networks for control systems: A survey
K. J. Hunt, D. Sbarbaro, R. Żbikowski, and P. J. Gawthrop · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. Williams · 1992
Earlier work this paper cites.
Active, uncalibrated visual servoing
B. H. Yoshimi and P. K. Allen · 1994
Earlier work this paper cites.
Skillful control under uncertainty via direct reinforcement learning
V. Gullapalli · 1995
Earlier work this paper cites.
Relative end-effector control using cartesian position based visual servoing
W. J. Wilson, C. W. Williams Hulls, and G. S. Bell · 1996
Earlier work this paper cites.
Biped dynamic walking using reinforcement learning
H. Benbrahim and J. A. Franklin · 1997
Earlier work this paper cites.
Experimental evaluation of uncalibrated visual servoing for precision manipulation
M. Jägersand, O. Fuentes, and R. C. Nelson · 1997
Earlier work this paper cites.
Neural Network Control of Robot Manipulators and Nonlinear Systems
F. L. Lewis, A. Yesildirak, and S. Jagannathan · 1998
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies
S. Hochreiter, Y. Bengio, P. Frasconi, and J. Schmidhuber · 2001
Earlier work this paper cites.
Covariant policy search
J. A. Bagnell and J. Schneider · 2003
Earlier work this paper cites.
A robot that reinforcement-learns to identify and memorize important previous observations
B. Bakker, V. Zhumatiy, G. Gruener, and J. Schmidhuber · 2003
Earlier work this paper cites.
Best practices for convolutional neural networks applied to visual document analysis
P. Y. Simard, D. Steinkraus, and J. C. Platt · 2003
Earlier work this paper cites.
Policy gradient reinforcement learning for fast quadrupedal locomotion
N. Kohl and P. Stone · 2004
Earlier work this paper cites.
Robotic surgery: a current perspective
A. Lanfranco, A. Castellanos, J. Desai, and W. Meyers · 2004
Earlier work this paper cites.
Iterative linear quadratic regulator design for nonlinear biological movement systems
W. Li and E. Todorov · 2004
Earlier work this paper cites.
Inverted autonomous helicopter flight via reinforcement learning
A. Y. Ng, H. J. Kim, M. I. Jordan, and S. Sastry · 2004
Earlier work this paper cites.
The Cross-Entropy Method: A Unified Approach to Combinatorial Optimization, Monte-Carlo Simulation and Machine Learning
R. Rubinstein and D. Kroese · 2004
Earlier work this paper cites.
Stochastic policy gradient reinforcement learning on a simple 3d biped
R. Tedrake, T. Zhang, and H. Seung · 2004
Earlier work this paper cites.
Learning to control an octopus arm with Gaussian process temporal difference methods
Y. Engel, P. Szabó, and D. Volkinshtein · 2005
Earlier work this paper cites.
Fast biped walking with a reflexive controller and realtime policy searching
T. Geng, B. Porr, and F. Wörgötter · 2006
Earlier work this paper cites.
A system for robotic heart surgery that learns to tie knots using recurrent neural networks
H. Mayer, F. Gomez, D. Wierstra, I. Nagy, A. Knoll, and J. Schmidhuber · 2006
Earlier work this paper cites.
Closed-loop learning of visual control policies
S. Jodogne and J. H. Piater · 2007
Earlier work this paper cites.
Applying the episodic natural actor-critic architecture to motor primitive learning
J. Peters and S. Schaal · 2007
Cited alongside, same era.
3D generic object categorization, localization and pose estimation
S. Savarese and L. Fei-Fei · 2007
Cited alongside, same era.
Learning CPG-based biped locomotion with a policy gradient method: Application to a humanoid robot
G. Endo, J. Morimoto, T. Matsubara, J. Nakanishi, and G. Cheng · 2008
Cited alongside, same era.
Reinforcement learning of motor skills with policy gradients
J. Peters and S. Schaal · 2008
Cited alongside, same era.
Towards a personal robotics development platform: Rationale and design of an intrinsically safe personal robot
K.A. Wyrobek, E.H. Berger, HF M. Van der Loos, and K. Salisbury · 2008
Cited alongside, same era.
ImageNet: A large-scale hierarchical image database
A survey on policy search for robotics
M. Deisenroth, G. Neumann, and J. Peters · 2013
Later among the works it cites.
Reinforcement learning in robotics: A survey
J. Kober, J. A. Bagnell, and J. Peters · 2013
Later among the works it cites.
Evolving large-scale neural networks for vision-based reinforcement learning
J. Koutník, G. Cuccu, J. Schmidhuber, and F. Gomez · 2013
Later among the works it cites.
Acquiring visual servoing reaching and grasping skills using neural reinforcement learning
T. Lampe and M. Riedmiller · 2013
Later among the works it cites.
Playing Atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Later among the works it cites.
Learning monocular reactive UAV control in cluttered natural environments
S. Ross, N. Melik-Barkhudarov, K. Shaurya Shankar, A. Wendel, D. Dey, J. A. Bagnell, and M. Hebert · 2013
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Deng, W. Dong, R. Socher, L. Li, K. Li, and L. Fei-Fei · 2009
Cited alongside, same era.
Learning long-range vision for autonomous off-road driving
R. Hadsell, P. Sermanet, J. B. A. Erkan, and M. Scoffier · 2009
Cited alongside, same era.
Learning motor primitives for robotics
J. Kober and J. Peters · 2009
Cited alongside, same era.
Convolutional deep belief networks for scalable unsupervised learning of hierarchical representations
H. Lee, R. Grosse, R. Ranganath, and A. Y. Ng · 2009
Cited alongside, same era.
Category independent object proposals
I. Endres and D. Hoiem · 2010
Cited alongside, same era.
BM: An iterative algorithm to learn stable non-linear dynamical systems with Gaussian mixture models
S. M. Khansari-Zadeh and A. Billard · 2010
Cited alongside, same era.
Autonomous door opening and plugging in with a personal robot
W. Meeussen, M. Wise, S. Glaser, S. Chitta, C. McGann, P. Mihelich, E. Marder-Eppstein, M. Muja, Victor Eruhimov, T. Foote, J. Hsu, R.B. Rusu, B. Marthi, G. Bradski, K. Konolige, B. Gerkey, and E. Berger · 2010
Cited alongside, same era.
Later among the works it cites.
Selective search for object recognition
J. Uijlings, K. van de Sande, T. Gevers, and A. Smeulders · 2013
Later among the works it cites.
Deep learning for real-time Atari game play using offline Monte-Carlo tree search planning
X. Guo, S. Singh, H. Lee, R. L. Lewis, and X. Wang · 2014
Later among the works it cites.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Later among the works it cites.
State representation learning in robotics: Using prior knowledge about physical interaction
R. Jonschkowski and O. Brock · 2014
Later among the works it cites.
Learning neural network policies with guided policy search under unknown dynamics
S. Levine and P. Abbeel · 2014
Later among the works it cites.
Learning complex neural network policies with trajectory optimization
S. Levine and V. Koltun · 2014
Later among the works it cites.
Sample-based information-theoretic stochastic optimal control
R. Lioutikov, A. Paraschos, G. Neumann, and J. Peters · 2014
Later among the works it cites.
Vision based control of a quadrotor for perching on planes and lines
K. Mohta, V. Kumar, and K. Daniilidis · 2014
Later among the works it cites.
Combining the benefits of function approximation and trajectory optimization
I. Mordatch and E. Todorov · 2014
Later among the works it cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2014
Later among the works it cites.
Joint training of a convolutional network and a graphical model for human pose estimation
J. J. Tompson, A. Jain, Y. LeCun, and C. Bregler · 2014
Later among the works it cites.
Bregman alternating direction method of multipliers
H. Wang and A. Banerjee · 2014
Later among the works it cites.
Learning visual feature spaces for robotic manipulation with deep spatial autoencoders
C. Finn, X. Tan, Y. Duan, T. Darrell, S. Levine, and P. Abbeel · 2015
Closest in time.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Closest in time.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Closest in time.
Learning contact-rich manipulation skills with guided policy search
S. Levine, N. Wagener, and P. Abbeel · 2015
Closest in time.
Continuous control with deep reinforcement learning
T. Lillicrap, J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Closest in time.
Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours
Lerrel Pinto and Abhinav Gupta · 2015
Closest in time.
Deep learning in neural networks: An overview
J. Schmidhuber · 2015
Closest in time.
Jaeyong Sung, Seok Hyun Jin, and Ashutosh Saxena · 2015
Closest in time.
Learning of non-parametric control policies with high-dimensional state features
H. van Hoof, J. Peters, and G. Neumann · 2015
Closest in time.
Embed to control: A locally linear latent dynamics model for control from raw images
M. Watter, J. Springenberg, J. Boedecker, and M. Riedmiller · 2015
Closest in time.