Fetching the paper…
Reading the bibliography…
We present an algorithm for rapidly learning controllers for robotics systems.
D. H. Jacobson and D. Q. Mayne, Differential Dynamic Programming . Elsevier, 1970
1970
Earlier work this paper cites.
C. G. Atkeson, A. W. Moore, and S. Schaal, “Locally weighted learning for control,” Lazy learning , pp. 75 – 113, 1997
1997
Earlier work this paper cites.
C. G. Atkeson and J. C. Santamaria, “A comparison of direct and model-based reinforcement learning,” in Proceedings of International Conference on Robotics and Automation , vol. 4, 1997
1997
Earlier work this paper cites.
N. L. Kleinman, J. C. Spall, and D. Q. Naiman, “Simulation-based optimization with stochastic approximation using common random numbers,” Management Science , 1999
1999
Earlier work this paper cites.
A. Y. Ng and M. Jordan, “Pegasus: A policy search method for large mdps and pomdps,” in Proceedings of the Sixteenth conference on Uncertainty in Artificial Intelligence , 2000
2000
Earlier work this paper cites.
W. Li and E. Todorov, “Iterative linear quadratic regulator design for nonlinear biological movement systems,” in Proceedings of the 1st International Conference on Informatics in Control, Automation and Robotics , 2004
2004
Earlier work this paper cites.
P. Abbeel, A. Coates, M. Quigley, and A. Y. Ng, “An application of reinforcement learning to aerobatic helicopter flight,” in Proceedings of Neural Information Processing Systems (NIPS) , 2006
2006
Earlier work this paper cites.
M. P. Deisenroth, Efficient reinforcement learning using Gaussian processes . KIT Scientific Publishing, 2010
2010
Earlier work this paper cites.
M. Lázaro-Gredilla, J. Quiñonero Candela, C. E. Rasmussen, and A. R. Figueiras-Vidal, “Sparse spectrum gaussian process regression,” Journal of Machine Learning Research , 2010
2010
Earlier work this paper cites.
D. Nguyen-Tuong and J. Peters, “Model learning for robot control: a survey,” Cognitive Processing , vol. 12, no. 4, pp. 319–340, Nov 2011
2011
Earlier work this paper cites.
R. Pascanu, T. Mikolov, and Y. Bengio, “On the difficulty of training recurrent neural networks,” in International Conference on Machine Learning , 2013
2013
Earlier work this paper cites.
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting.” Journal of Machine Learning Research , 2014
2014
Cited alongside, same era.
Y. Tassa, N. Mansard, and E. Todorov, “Control-limited differential dynamic programming,” in Proceedings of the International Conference on Robotics and Automation (ICRA) , 2014
2014
Cited alongside, same era.
S. Levine and P. Abbeel, “Learning neural network policies with guided policy search under unknown dynamics,” in Advances in Neural Information Processing Systems , 2014, pp. 1071–1079
2014
Cited alongside, same era.
Y. Pan and E. Theodorou, “Probabilistic differential dynamic programming,” in Advances in Neural Information Processing Systems , 2014
2014
Cited alongside, same era.
D. P. Kingma, T. Salimans, and M. Welling, “Variational dropout and the local reparameterization trick,” in Advances in Neural Information Processing Systems , 2015
2015
Later among the works it cites.
Y. Gal, R. McAllister, and C. E. Rasmussen, “Improving PILCO with Bayesian neural network dynamics models,” in Data-Efficient Machine Learning workshop, ICML , 2016
2016
Later among the works it cites.
Y. Gal and Z. Ghahramani, “Dropout as a bayesian approximation: Representing model uncertainty in deep learning,” in Proceedings of The 33rd International Conference on Machine Learning , 2016
2016
Later among the works it cites.
S. Gu, T. Lillicrap, I. Sutskever, and S. Levine, “Continuous deep q-learning with model-based acceleration,” in International Conference on Machine Learning , 2016
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
M. P. Deisenroth, D. Fox, and C. E. Rasmussen, “Gaussian processes for data-efficient learning in robotics and control,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2015
2015
Cited alongside, same era.
D. Meger, J. C. G. Higuera, A. Xu, P. Giguere, and G. Dudek, “Learning legged swimming gaits from experience,” in Robotics and Automation (ICRA), 2015 IEEE International Conference on , 2015
2015
Cited alongside, same era.
A. D. Marchese, R. Tedrake, and D. Rus, “Dynamics and trajectory optimization for a soft spatial fluidic elastomer manipulator,” International Journal of Robotics Research , 2015
2015
Cited alongside, same era.
N. Heess, G. Wayne, D. Silver, T. Lillicrap, T. Erez, and Y. Tassa, “Learning continuous control policies by stochastic value gradients,” in Advances in Neural Information Processing Systems , 2015
2015
Cited alongside, same era.
2015
Cited alongside, same era.
C. Blundell, J. Cornebise, K. Kavukcuoglu, and D. Wierstra, “Weight uncertainty in neural network,” in International Conference on Machine Learning , 2015
2015
Cited alongside, same era.
K. Neklyudov, D. Molchanov, A. Ashukha, and D. P. Vetrov, “Structured bayesian pruning via log-normal multiplicative noise,” in Advances in Neural Information Processing Systems , 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
K. Chatzilygeroudis, R. Rama, R. Kaushik, D. Goepp, V. Vassiliades, and J.-B. Mouret, “Black-box data-efficient policy search for robotics,” in Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , 2017
2017
Later among the works it cites.
D. Molchanov, A. Ashukha, and D. Vetrov, “Variational dropout sparsifies deep neural networks,” in International Conference on Machine Learning , 2017
2017
Later among the works it cites.
Y. Gal, J. Hron, and A. Kendall, “Concrete dropout,” in Advances in Neural Information Processing Systems , 2017
2017
Later among the works it cites.