Fetching the paper…
Reading the bibliography…
Trust-region methods have yielded state-of-the-art results in policy search.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Natural gradient works efficiently in learning
S. Amari · 1998
Earlier work this paper cites.
The cross-entropy method for combinatorial and continuous optimization
Reuven Rubinstein · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Completely derandomized self-adaptation in evolution strategies
Nikolaus Hansen and Andreas Ostermeier · 2001
Earlier work this paper cites.
A natural policy gradient
Sham Kakade · 2001
Earlier work this paper cites.
Covariant policy search
J Andrew Bagnell and Jeff Schneider · 2003
Earlier work this paper cites.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Natural actor-critic
J. Peters and S. Schaal · 2008
Earlier work this paper cites.
Online planning algorithms for POMDPs
Stéphane Ross, Joelle Pineau, Sébastien Paquet, and Brahim Chaib-Draa · 2008
Earlier work this paper cites.
Natural evolution strategies
D. Wierstra, T. Schaul, J. Peters, and J. Schmidhuber · 2008
Earlier work this paper cites.
Policy search for motor primitives in robotics
Jens Kober and Jan R. Peters · 2009
Cited alongside, same era.
Revisiting natural actor-critics with value function approximation
Matthieu Geist and Olivier Pietquin · 2010
Cited alongside, same era.
Relative entropy policy search
Jan Peters, Katharina Mülling, and Yasemin Altun · 2010
Cited alongside, same era.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
Model-based relative entropy stochastic search
A. Abdolmaleki, R. Lioutikov, J Peters, N. Lau, L. Reis, and G. Neumann · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Later among the works it cites.
PGQ: combining policy gradient and q-learning
Brendan O’Donoghue, Rémi Munos, Koray Kavukcuoglu, and Volodymyr Mnih · 2016
Later among the works it cites.
CARLA: An Open Urban Driving Simulator
Alexey Dosovitskiy, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun · 2017
Later among the works it cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Model-free trajectory optimization for reinforcement learning
R. Akrour, A. Abdolmaleki, H. Abdulsamad, and G. Neumann · 2016
Cited alongside, same era.
Hierarchical relative entropy policy search
C. Daniel, G. Neumann, O. Kroemer, and J. Peters · 2016
Cited alongside, same era.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Yuhuai Wu, Elman Mansimov, Roger B Grosse, Shun Liao, and Jimmy Ba · 2017
Later among the works it cites.
Maximum a Posteriori Policy Optimisation
Abbas Abdolmaleki, Jost T. Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, and Martin Riedmiller · 2018
Later among the works it cites.
Model-Free Trajectory-based Policy Optimization with Monotonic Improvement
Riad Akrour, Abbas Abdolmaleki, Hany Abdulsamad, Jan Peters, and Gerhard Neumann · 2018
Later among the works it cites.
Exact natural gradient in deep linear networks and its application to the nonlinear case
Alberto Bernacchia, Mate Lengyel, and Guillaume Hennequin · 2018
Later among the works it cites.
Guide Actor-Critic for Continuous Control
Voot Tangkaratt, Abbas Abdolmaleki, and Masashi Sugiyama · 2018
Later among the works it cites.