Fetching the paper…
Reading the bibliography…
Trust region methods, such as TRPO, are often used to stabilize policy optimization algorithms in reinforcement learning (RL).
Learning from delayed rewards
Christopher John Cornish Hellaby Watkins · 1989
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Ronald J Williams and Jing Peng · 1991
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Actor-critic algorithms, 2000
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2002
Earlier work this paper cites.
Covariant policy search
J Andrew Bagnell and Jeff Schneider · 2003
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
Jan Peters and Stefan Schaal · 2008
Earlier work this paper cites.
Relative entropy policy search
Jan Peters, Katharina Mulling, and Yasemin Altun · 2010
Earlier work this paper cites.
Dynamic policy programming with function approximation
Mohammad Gheshlaghi Azar, Vicenç Gómez, and Hilbert J Kappen · 2011
Earlier work this paper cites.
Differential-geometrical methods in statistics , volume 28
Shun-Ichi Amari · 2012
Earlier work this paper cites.
Dynamic policy programming
Mohammad Gheshlaghi Azar, Vicenç Gómez, and Hilbert J Kappen · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller · 2013
Cited alongside, same era.
On stochastic optimal control and reinforcement learning by approximate inference
Konrad Rawlik, Marc Toussaint, and Sethu Vijayakumar · 2013
Cited alongside, same era.
Proximal algorithms
Neal Parikh, Stephen Boyd, et al · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Vime: Variational information maximizing exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy P Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Reward augmented maximum likelihood for neural structured prediction
Mohammad Norouzi, Samy Bengio, Zhifeng Chen, Navdeep Jaitly, Mike Schuster, Yonghui Wu, and Dale Schuurmans · 2016
Later among the works it cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2016
Later among the works it cites.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Cited alongside, same era.
G-learning: Taming the noise in reinforcement learning via soft updates
Roy Fox, Ari Pakman, and Naftali Tishby · 2016
Cited alongside, same era.
Ziyu Wang, Victor Bapst, Nicolas Heess, Volodymyr Mnih, Remi Munos, Koray Kavukcuoglu, and Nando de Freitas · 2016
Later among the works it cites.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Closest in time.
Emergence of locomotion behaviours in rich environments
Nicolas Heess, Srinivasan Sriram, Jay Lemmon, Josh Merel, Greg Wayne, Yuval Tassa, Tom Erez, Ziyu Wang, Ali Eslami, Martin Riedmiller, et al · 2017
Closest in time.
Discrete sequential prediction of continuous actions for deep RL
Luke Metz, Julian Ibarz, Navdeep Jaitly, and James Davidson · 2017
Closest in time.
Bridging the gap between value and policy based reinforcement learning
Ofir Nachum, Mohammad Norouzi, Kelvin Xu, and Dale Schuurmans · 2017
Closest in time.
Modular rl
John Schulman · 2017
Closest in time.