Fetching the paper…
Reading the bibliography…
Policy gradients methods apply to complex, poorly understood, control problems by performing stochastic gradient descent over a parameterized class of polices.
Contributions to the theory of optimal control
Rudolf Emil Kalman et al · 1960
Earlier work this paper cites.
Gradient methods for the minimisation of functionals
Boris T Polyak · 1963
Earlier work this paper cites.
Discounted dynamic programming
David Blackwell · 1965
Earlier work this paper cites.
On an iterative technique for riccati equation computations
D. Kleinman · 1968
Earlier work this paper cites.
An iterative technique for the computation of the steady state gains for the discrete optimal regulator
G Hewer · 1971
Earlier work this paper cites.
Stochastic optimal control: the discrete-time case
Dimitir P Bertsekas and Steven Shreve · 1978
Earlier work this paper cites.
A globally convergent algorithm for the optimal constant output feedback problem
Hannu T Toivonen · 1985
Earlier work this paper cites.
The mathematics of nonlinear programming
Anthony L Peressini, Francis E Sullivan, and J Jerry Uhl · 1988
Earlier work this paper cites.
Derivatives of probability measures-concepts and applications to the optimization of stochastic systems
G. Ch. Pflug · 1988
Earlier work this paper cites.
On-line optimization of simulated markovian processes
G. Ch. Pflug · 1990
Earlier work this paper cites.
Gradient estimation via perturbation analysis , volume 116
Paul Glasserman and Yu-Chi Ho · 1991
Earlier work this paper cites.
Perturbation analysis for the design of flexible manufacturing system flow controllers
Michael Caramanis and George Liberopoulos · 1992
Earlier work this paper cites.
Efficient exploration in reinforcement learning
Sebastian B Thrun · 1992
Earlier work this paper cites.
Complexity analysis of real-time reinforcement learning
Sven Koenig and Reid G Simmons · 1993
Earlier work this paper cites.
Inequalities for the trace of matrix product
Yuguang Fang, Kenneth A Loparo, and Xiangbo Feng · 1994
Earlier work this paper cites.
Stochastic optimization by simulation: Convergence proofs for the gi/g/1 queue in steady-state
Pierre L’Ecuyer and Peter W Glynn · 1994
Earlier work this paper cites.
Stochastic optimization by simulation: numerical experiments with the m/m/1 queue in steady-state
Pierre L’Ecuyer, Nataly Giroux, and Peter W Glynn · 1994
Earlier work this paper cites.
Reinforcement learning with soft state aggregation
Satinder P Singh, Tommi Jaakkola, and Michael I Jordan · 1994
Earlier work this paper cites.
Dynamic programming and optimal control
Dimitri P Bertsekas · 1995
Earlier work this paper cites.
Sensitivity analysis for base-stock levels in multiechelon production-inventory systems
Paul Glasserman and Sridhar Tayur · 1995
Earlier work this paper cites.
Neuro-dynamic programming , volume 5
Dimitri P Bertsekas and John N Tsitsiklis · 1996
Earlier work this paper cites.
Feature-based methods for large scale dynamic programming
John N Tsitsiklis and Benjamin Van Roy · 1996
Earlier work this paper cites.
Nonlinear programming
Dimitri P Bertsekas · 1997
Earlier work this paper cites.
Computational design of optimal output feedback controllers
Tankred Rautert and Ekkehard W Sachs · 1997
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Jonathan Baxter and Peter L Bartlett · 2001
Earlier work this paper cites.
Simulation-based optimization of markov reward processes
Peter Marbach and John N Tsitsiklis · 2001
Earlier work this paper cites.
Convergence rate analysis of gradient based algorithms
Amir Beck · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2002
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Amir Beck and Marc Teboulle · 2003
Earlier work this paper cites.
Error bounds for approximate policy iteration
Rémi Munos · 2003
Earlier work this paper cites.
Policy search by dynamic programming
Andrew Bagnell, Sham M Kakade, Jeff G Schneider, and Andrew Y Ng · 2004
Earlier work this paper cites.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Metrics for finite markov decision processes
Norm Ferns, Prakash Panangaden, and Doina Precup · 2004
Earlier work this paper cites.
An introduction to mathematical optimal control theory
Lawrence C Evans · 2005
Earlier work this paper cites.
Error bounds for approximate value iteration
Rémi Munos · 2005
Cited alongside, same era.
Gradient estimation
Michael C Fu · 2006
Cited alongside, same era.
Cubic regularization of newton method and its global performance
Yurii Nesterov and Boris T Polyak · 2006
Cited alongside, same era.
Policy gradient methods for robotics
Jan Peters and Stefan Schaal · 2006
Cited alongside, same era.
The theory and practice of revenue management , volume 68
Kalyan T Talluri and Garrett J Van Ryzin · 2006
Cited alongside, same era.
Performance loss bounds for approximate value iteration with state aggregation
Benjamin Van Roy · 2006
Cited alongside, same era.
Stochastic simulation: algorithms and analysis , volume 57
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2016
Later among the works it cites.
Finding approximate local minima faster than gradient descent
Naman Agarwal, Zeyuan Allen-Zhu, Brian Bullins, Elad Hazan, and Tengyu Ma · 2017
Later among the works it cites.
First-order methods in optimization , volume 25
Amir Beck · 2017
Later among the works it cites.
Is the bellman residual a bad proxy?
Matthieu Geist, Bilal Piot, and Olivier Pietquin · 2017
Later among the works it cites.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M Kakade, and Michael I Jordan · 2017
Later among the works it cites.
Towards generalization and simplicity in continuous control
Aravind Rajeswaran, Kendall Lowrey, Emanuel V Todorov, and Sham M Kakade · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Søren Asmussen and Peter W Glynn · 2007
Cited alongside, same era.
Performance bounds in L_p-norm for approximate value iteration
Rémi Munos · 2007
Cited alongside, same era.
Evaluation of policy gradient methods and variants on the cart-pole benchmark
Martin Riedmiller, Jan Peters, and Stefan Schaal · 2007
Cited alongside, same era.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Cited alongside, same era.
Using stochastic approximation methods to compute optimal base-stock levels in inventory control problems
Sumit Kunnumkal and Huseyin Topaloglu · 2008
Cited alongside, same era.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Cited alongside, same era.
Chang-Han Rhee and Peter Glynn · 2017
Later among the works it cites.
Evolution strategies as a scalable alternative to reinforcement learning
Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever · 2017
Later among the works it cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Complete dictionary recovery over the sphere I: Overview and the geometric picture
J. Sun, Q. Qu, and J. Wright · 2017
Later among the works it cites.
Accelerated methods for nonconvex optimization
Yair Carmon, John C Duchi, Oliver Hinder, and Aaron Sidford · 2018
Later among the works it cites.
On the power of over-parametrization in neural networks with quadratic activation
Simon S Du and Jason D Lee · 2018
Later among the works it cites.
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham Kakade, and Mehran Mesbahi · 2018
Later among the works it cites.
Variational inverse control with events: A general framework for data-driven reward definition
Justin Fu, Avi Singh, Dibya Ghosh, Larry Yang, and Sergey Levine · 2018
Later among the works it cites.
Soft actor-critic algorithms and applications
Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, et al · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Feature-based aggregation and deep reinforcement learning: A survey and some new implementations
Dimitri P Bertsekas · 2019
Closest in time.
Lqr through the lens of first order methods: Discrete-time case
Jingjing Bu, Afshin Mesbahi, Maryam Fazel, and Mehran Mesbahi · 2019
Closest in time.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Closest in time.
Proximally guided stochastic subgradient method for nonsmooth, nonconvex problems
Damek Davis and Benjamin Grimmer · 2019
Closest in time.
A theory of regularized markov decision processes
Matthieu Geist, Bruno Scherrer, and Olivier Pietquin · 2019
Closest in time.
Deep exploration via randomized value functions
Ian Osband, Benjamin Van Roy, Daniel J Russo, and Zheng Wen · 2019
Closest in time.
Neural policy gradient methods: Global optimality and rates of convergence
Lingxiao Wang, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2019
Closest in time.
Optimality and approximation with policy gradient methods in markov decision processes
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2020
Closest in time.
Stochastic subgradient method converges on tame functions
Damek Davis, Dmitriy Drusvyatskiy, Sham Kakade, and Jason D Lee · 2020
Closest in time.
Is a good representation sufficient for sample efficient reinforcement learning?
Simon S. Du, Sham M. Kakade, Ruosong Wang, and Lin F. Yang · 2020
Closest in time.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Closest in time.
Continual learning for robotics: Definition, framework, learning strategies, opportunities and challenges
Timothée Lesort, Vincenzo Lomonaco, Andrei Stoian, Davide Maltoni, David Filliat, and Natalia Díaz-Rodríguez · 2020
Closest in time.
Monte carlo gradient estimation in machine learning
Shakir Mohamed, Mihaela Rosca, Michael Figurnov, and Andriy Mnih · 2020
Closest in time.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps
Lior Shani, Yonathan Efroni, and Shie Mannor · 2020
Closest in time.
The ingredients of real-world robotic reinforcement learning
Henry Zhu, Justin Yu, Abhishek Gupta, Dhruv Shah, Kristian Hartikainen, Avi Singh, Vikash Kumar, and Sergey Levine · 2020
Closest in time.
Global optimiality guarantees for policy gradient methods: Technical report with supplementary materials
Jalaj Bhandari and Daniel Russo · 2021
Closest in time.
Masatoshi Uehara, Masaaki Imaizumi, Nan Jiang, Nathan Kallus, Wen Sun, and Tengyang Xie · 2021
Closest in time.
What are the statistical limits of offline RL with linear function approximation?
Ruosong Wang, Dean Foster, and Sham M. Kakade · 2021
Closest in time.
Exponential lower bounds for planning in mdps with linearly-realizable optimal action-value functions
Gellért Weisz, Philip Amortila, and Csaba Szepesvári · 2021
Closest in time.