Fetching the paper…
Reading the bibliography…
Policy gradient methods are among the most effective methods in challenging reinforcement learning problems with large state and/or action spaces.
Neural temporal-difference learning converges to global optima
Qi Cai, Zhuoran Yang, Jason D. Lee, and Zhaoran Wang · 1905
Earlier work this paper cites.
Global optimality guarantees for policy gradient methods
Jalaj Bhandari and Daniel Russo · 1906
Earlier work this paper cites.
Neural proximal/trust region policy optimization attains globally optimal policy
Boyi Liu, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 1906
Earlier work this paper cites.
Yasin Abbasi-Yadkori, Nevena Lazic, Csaba Szepesvari, and Gellert Weisz · 1908
Earlier work this paper cites.
Extremum problems with inequalities as subsidiary conditions
Fritz John · 1948
Earlier work this paper cites.
Functional approximations and dynamic programming
Richard Bellman and Stuart Dreyfus · 1959
Earlier work this paper cites.
Gradient methods for minimizing functionals
B. T. Polyak · 1963
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Arkadii Semenovich Nemirovsky and David Borisovich Yudin · 1983
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Ronald J Williams and Jing Peng · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Neuro-Dynamic Programming
Dimitri P Bertsekas and John N Tsitsiklis · 1996
Earlier work this paper cites.
An elementary introduction to modern convex geometry
Keith Ball · 1997
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E Schapire · 1997
Earlier work this paper cites.
Least-squares temporal difference learning
Justin A Boyan · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
A natural policy gradient
S. Kakade · 2001
Earlier work this paper cites.
Approximately Optimal Approximate Reinforcement Learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Covariant policy search
J. Andrew Bagnell and Jeff Schneider · 2003
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2003
Earlier work this paper cites.
Error bounds for approximate policy iteration
Rémi Munos · 2003
Earlier work this paper cites.
Policy search by dynamic programming
J. A. Bagnell, Sham M Kakade, Jeff G. Schneider, and Andrew Y. Ng · 2004
Cited alongside, same era.
Error bounds for approximate value iteration
Rémi Munos · 2005
Cited alongside, same era.
Finite time bounds for sampling based fitted value iteration
Csaba Szepesvári and Rémi Munos · 2005
Cited alongside, same era.
Prediction, Learning, and Games
Nicolo Cesa-Bianchi and Gabor Lugosi · 2006
Cited alongside, same era.
Cubic regularization of newton method and its global performance
Yurii Nesterov and Boris T. Polyak · 2006
Cited alongside, same era.
The łojasiewicz inequality for nonsmooth subanalytic functions with applications to subgradient dynamical systems
Jérôme Bolte, Aris Daniilidis, and Adrian Lewis · 2007
Cited alongside, same era.
Escaping from saddle points - online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Later among the works it cites.
Approximate modified policy iteration and its application to the game of tetris
Bruno Scherrer, Mohammad Ghavamzadeh, Victor Gabillon, Boris Lesner, and Matthieu Geist · 2015
Later among the works it cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Later among the works it cites.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
Saeed Ghadimi and Guanghui Lan · 2016
Later among the works it cites.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Later among the works it cites.
Analysis of classification-based policy iteration algorithms
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Cited alongside, same era.
Natural actor-critic
Jan Peters and Stefan Schaal · 2008
Cited alongside, same era.
Natural actor–critic algorithms
Shalabh Bhatnagar, Richard S Sutton, Mohammad Ghavamzadeh, and Mark Lee · 2009
Cited alongside, same era.
Online Markov decision processes
Eyal Even-Dar, Sham M Kakade, and Yishay Mansour · 2009
Cited alongside, same era.
Proximal alternating minimization and projection methods for nonconvex problems: An approach based on the kurdyka-łojasiewicz inequality
Hédy Attouch, Jérôme Bolte, Patrick Redont, and Antoine Soubeyran · 2010
Cited alongside, same era.
Error propagation for approximate policy and value iteration
Amir-massoud Farahmand, Csaba Szepesvári, and Rémi Munos · 2010
Cited alongside, same era.
Alessandro Lazaric, Mohammad Ghavamzadeh, and Rémi Munos · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Later among the works it cites.
First-Order Methods in Optimization
A. Beck · 2017
Later among the works it cites.
Contextual decision processes with low bellman rank are PAC-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Later among the works it cites.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M. Kakade, and Michael I. Jordan · 2017
Later among the works it cites.
A unified view of entropy-regularized markov decision processes
Gergely Neu, Anders Jonsson, and Vicenç Gómez · 2017
Later among the works it cites.
Towards generalization and simplicity in continuous control
Aravind Rajeswaran, Kendall Lowrey, Emanuel V. Todorov, and Sham M Kakade · 2017
Later among the works it cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Maximum a posteriori policy optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, and Martin Riedmiller · 2018
Later among the works it cites.
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham M Kakade, and Mehran Mesbahi · 2018
Later among the works it cites.
Understanding the impact of entropy on policy optimization , 2019
Zafarali Ahmed, Nicolas Le Roux, Mohammad Norouzi, and Dale Schuurmans, editors · 2019
Closest in time.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Closest in time.
A theory of regularized markov decision processes
Matthieu Geist, Bruno Scherrer, and Olivier Pietquin · 2019
Closest in time.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2019
Closest in time.
Adaptive trust region policy optimization: Global convergence and fa ster rates for regularized mdps, 2019
Lior Shani, Yonathan Efroni, and Shie Mannor · 2019
Closest in time.
Sample-optimal parametric q-learning using linearly additive features
Lin F. Yang and Mengdi Wang · 2019
Closest in time.