Fetching the paper…
Reading the bibliography…
Direct policy gradient methods for reinforcement learning and continuous control problems are a popular approach for a variety of reasons: 1) they are easy to implement without explicit knowledge of the underlying model 2) they are an "end-to-end" approach, directly optimizing the performance metric of interest 3) they inherently allow for richly parameterized policies.
A survey on policy search for robotics
Marc Peter Deisenroth, Gerhard Neumann, and Jan Peters · 1935
Earlier work this paper cites.
Gradient methods for minimizing functionals
B. T. Polyak · 1963
Earlier work this paper cites.
Dynamic programming and Markov processes
Ronald A Howard · 1964
Earlier work this paper cites.
On an iterative technique for Riccati equation computations
D. L. Kleinman · 1968
Earlier work this paper cites.
An iterative technique for the computation of steady state gains for the discrete optimal regulator
G. A. Hewer · 1971
Earlier work this paper cites.
An Historical Survey of Computational Methods in Optimal Control
E. Polak · 1973
Earlier work this paper cites.
Optimal Control: Linear Quadratic Methods
Brian D. O. Anderson and John B. Moore · 1990
Earlier work this paper cites.
Matrix perturbation theory (computer science and scientific computing), 1990
Gilbert W Stewart and Ji-Guang Sun · 1990
Earlier work this paper cites.
Convergence in unconstrained discrete-time differential dynamic programming
L. Z. Liao and C. A. Shoemaker · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Adaptive linear quadratic control using policy iteration
S.J. Bradtke, B.E. Ydstie, and a.G. Barto · 1994
Earlier work this paper cites.
PAC adaptive control of liner systems
Claude-Nicolas Fiechter · 1994
Earlier work this paper cites.
Algebraic Riccati Equations
P. Lancaster and L. Rodman · 1995
Earlier work this paper cites.
System Identification (2Nd Ed.): Theory for the User
Lennart Ljung, editor · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
A natural policy gradient
S. Kakade · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
S. Kakade and J. Langford · 2002
Earlier work this paper cites.
A globally convergent and efficient method for unconstrained discrete-time optimal control
Chi-Kong Ng, Li-Zhi Liao, and Duan Li · 2002
Earlier work this paper cites.
Covariant policy search
J. Andrew Bagnell and Jeff Schneider · 2003
Earlier work this paper cites.
Semidefinite programming duality and linear time-invariant systems
V. Balakrishnan and L. Vandenberghe · 2003
Cited alongside, same era.
On the sample complexity of reinforcement learning
S. M. Kakade · 2003
Cited alongside, same era.
Model Predictive Control
E.F. Camacho and C. Bordons · 2004
Cited alongside, same era.
A generalized iterative lqg method for locally-optimal feedback control of constrained nonlinear stochastic systems
Emanuel Todorov and Weiwei Li · 2004
Cited alongside, same era.
An introduction to mathematical optimal control theory
Lawrence C. Evans · 2005
Cited alongside, same era.
Online convex optimization in the bandit setting: gradient descent without a gradient
Abraham D Flaxman, Adam Tauman Kalai, and H Brendan McMahan · 2005
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Later among the works it cites.
Random gradient-free minimization of convex functions
Yurii Nesterov and Vladimir Spokoiny · 2015
Later among the works it cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Moritz, M. Jordan, and P. Abbeel · 2015
Later among the works it cites.
Benchmarking deep reinforcement learning for continuous control
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel · 2016
Later among the works it cites.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cubic regularization of newton method and its global performance
Yurii Nesterov and Boris T. Polyak · 2006
Cited alongside, same era.
Natural actor-critic
J. Peters and S. Schaal · 2007
Cited alongside, same era.
Introduction to derivative-free optimization , volume 8 of MPS/SIAM Series on Optimization
A.R. Conn, K. Scheinberg, and L.N. Vicente · 2009
Cited alongside, same era.
Gradient methods for iterative distributed control synthesis
Karl Mårtensson and Anders Rantzer · 2009
Cited alongside, same era.
Model Predictive Control: Theory and Design
J.B. Rawlings and D.Q. Mayne · 2009
Cited alongside, same era.
Regret bounds for the adaptive control of linear quadratic systems
Yasin Abbasi-Yadkori and Csaba Szepesvári · 2011
Cited alongside, same era.
Optimal control with learned local models: Application to dexterous manipulation
V. Kumar, E. Todorov, and S. Levine · 2016
Later among the works it cites.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Later among the works it cites.
Dynamic Programming and Optimal Control
Dimitri P. Bertsekas · 2017
Later among the works it cites.
On the sample complexity of the linear quadratic regulator
S. Dean, H. Mania, N. Matni, B. Recht, and S. Tu · 2017
Later among the works it cites.
Learning linear dynamical systems via spectral filtering
Elad Hazan, Karan Singh, and Cyril Zhang · 2017
Later among the works it cites.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M. Kakade, and Michael I. Jordan · 2017
Later among the works it cites.
Evolution strategies as a scalable alternative to reinforcement learning
Tim Salimans, Jonathan Ho, Xi Chen, and Ilya Sutskever · 2017
Later among the works it cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel · 2017
Later among the works it cites.
Towards provable control for unknown linear dynamical systems
Sanjeev Arora, Elad Hazan, Holden Lee, Karan Singh, Cyril Zhang, and Yi Zhang · 2018
Closest in time.
Spectral filtering for general linear dynamical systems
Elad Hazan, Holden Lee, Karan Singh, Cyril Zhang, and Yi Zhang · 2018
Closest in time.
Learning without mixing: Towards a sharp analysis of linear system identification
Max Simchowitz, Horia Mania, Stephen Tu, Michael I Jordan, and Benjamin Recht · 2018
Closest in time.
Least-squares temporal difference learning for the linear quadratic regulator
Stephen Tu and Benjamin Recht · 2018
Closest in time.