Fetching the paper…
Reading the bibliography…
Recent advances in policy gradient methods and deep learning have demonstrated their applicability for complex reinforcement learning problems.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning , vol. 8, p. 229–256, 1992
1992
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Introduction to Reinforcement Learning . Cambridge, MA, USA: MIT Press, 1998
1998
Earlier work this paper cites.
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” in NIPS’99 Proceedings of the 12th International Conference on Neural Information Processing Systems , 1999, pp. 1057–1063
1999
Earlier work this paper cites.
D. P. Bertsekas, Nonlinear Programming . Belmont, MA, USA: Athena Scientific, 1999
1999
Earlier work this paper cites.
A. Owen and Y. Zhou, “Safe and effective importance sampling,” Journal of the American Statistical Association , vol. 95, pp. 135–143, 2000
2000
Earlier work this paper cites.
S. Kakade, “A natural policy gradient,” in Proceedings of the 14th International Conference on Neural Information Processing Systems (NIPS-01) . MIT Press, 2001, pp. 1531–1538
2001
Earlier work this paper cites.
N. Kohl and P. Stone, “Policy gradient reinforcement learning for fast quadrupedal locomotion,” in Proceedings of the IEEE International Conference on Robotics and Automation (ICRA 2004) , 2004
2004
Earlier work this paper cites.
R. Tedrake, T. W. Zhang, and H. S. Seung, “Stochastic policy gradient reinforcement learning on a simple 3d biped,” in IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2004) , 2004
2004
Earlier work this paper cites.
J. Nocedal and S. J. Wright, Numerical Optimization . New York, NY, USA: Springer, 2006
2006
Earlier work this paper cites.
J. Peters and S. Schaal, “Reinforcement learning of motor skills with policy gradients,” Neural Networks , vol. 21, pp. 682–697, 2008
2008
Earlier work this paper cites.
J. Peters and S. Schaal, “Natural actor-critic,” Neurocomputing , vol. 71, pp. 1180–1190, 2008
2008
Earlier work this paper cites.
R. H. Byrd, G. M. Chin, W. Neveittand, and J. Nocedal, “On the use of stochastic hessian information in optimization methods for machine learning,” SIAM Journal on Optimization (SIOPT) , vol. 21, p. 977–995, 2011
2011
Cited alongside, same era.
M. P. Deisenroth and C. E. Rasmussen, “Pilco: A model-based and data-efficient approach to policy search,” in Proceedings of the 28th International Conference on Machine Learning (ICML 2011) , 2011
2011
Cited alongside, same era.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems , pp. 5026–5033, 2012
2012
Cited alongside, same era.
R. Johnson and T. Zhang, “Accelerating stochastic gradient descent using predictive variance reduction,” in Proceedings of the 26th International Conference on Neural Information Processing Systems (NIPS) , 2013
2013
Cited alongside, same era.
S. Levine, N. Wagener, and P. Abbeel, “Learning contact-rich manipulation skills with guided policy search,” in Proceedings of the IEEE International Conference on Robotics and Automation (ICRA 2015) , 2015
2015
Later among the works it cites.
2016
Later among the works it cites.
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel, “Benchmarking deep reinforcement learning for continuous control,” in Proceedings of the 33rd International Conference on Machine Learning (ICML) , 2016
2016
Later among the works it cites.
S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,” The Journal of Machine Learning Research , vol. 17, pp. 1334–1373, 2016
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. P. Deisenroth, G. Neumann, and J. Peters, “A survey on policy search for robotics,” Foundations and Trends in Robotics , vol. 2, pp. 1–142, 2013
2013
Cited alongside, same era.
M. P. Deisenroth, P. Englert, J. Peters, and D. Fox, “Multi-task policy search for robotics,” in Proceedings of the IEEE International Conference on Robotics and Automation (ICRA 2014) , 2014
2014
Cited alongside, same era.
B. Bischoff, D. Nguyen-Tuong, H. van Hoof, A. McHutchon, C. Rasmussen, A. Knoll, J. Peters, and M. Deisenroth, “Policy search for learning robot control using sparse data,” in Proceedings of the IEEE International Conference on Robotics and Automation (ICRA 2014) , 2014, pp. 3882–3887
2014
Cited alongside, same era.
Y. Nesterov, Introductory Lectures on Convex Optimization . Springer, 2014
2014
Cited alongside, same era.
J. Schulman, S. Levine, P. Moritz, M. I. Jordan, and P. Abbeel, “Trust region policy optimization,” in Proceedings of the 32nd International Conference on Machine Learning (ICML) , 2015
2015
Cited alongside, same era.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, “Human-level control through deep reinforcement learning,” Nature , vol. 518, pp. 529–533, 2015
2015
Cited alongside, same era.
R. Kolte, M. Erdogdu, and A. Ozgür, “Accelerating svrg via second-order information,” In NIPS Workshop on Optimization for Machine Learning , 2015
2015
Cited alongside, same era.
R. N. P. Moritz and M. Jordan, “A linearly-convergent stochastic l-bfgs algorithm,” in Proceedings of the 19th International Conference on Artificial Intelligence and Statistics (AISTATS) , 2016, pp. 249–258
2016
Later among the works it cites.
2016
Later among the works it cites.
2017
Closest in time.
M. Sheckells, G. Garimella, and M. Kobilarov, “Robust policy search with applications to safe vehicle navigation,” in Proceedings of the IEEE International Conference on Robotics and Automation (ICRA 2017) , 2017, pp. 2343–2349
2017
Closest in time.
C. Amato, G. Konidaris, A. Anders, G. Cruz, J. P. How, and L. P. Kaelbling, “Policy search for multi-robot coordination under uncertainty,” The International Journal of Robotics Research , vol. 35, 2017
2017
Closest in time.
S. S. Du, J. Chen, L. Li, L. Xiao, and D. Zhou, “Stochastic variance reduction methods for policy evaluation,” in Proceedings of the International Conference on Machine Learning(ICML) , 2017
2017
Closest in time.