Fetching the paper…
Reading the bibliography…
We consider infinite-horizon discounted Markov decision processes and study the convergence rates of the natural policy gradient (NPG) and the Q-NPG methods with the log-linear policy class.
Une propriété topologique des sous-ensembles analytiques réels
Stanisław Łojasiewicz · 1963
Earlier work this paper cites.
Gradient methods for the minimisation of functionals
B.T. Polyak · 1963
Earlier work this paper cites.
The relaxation method of finding the common point of convex sets and its application to the solution of problems in convex programming
L.M. Bregman · 1967
Earlier work this paper cites.
Convex analysis
R. Tyrrell Rockafellar · 1970
Earlier work this paper cites.
Problem Complexity and Method Efficiency in Optimization
Arkadi Nemirovski and David Berkovich Yudin · 1983
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Convergence analysis of a proximal-like minimization algorithm using bregman functions
Gong Chen and Marc Teboulle · 1993
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
John Tsitsiklis and Benjamin Van Roy · 1996
Earlier work this paper cites.
Parallel Optimization: Theory, Algorithms, and Applications
Y. Censor and S.A. Zenios · 1997
Earlier work this paper cites.
Natural Gradient Works Efficiently in Learning
Shun-ichi Amari · 1998
Earlier work this paper cites.
Actor-critic algorithms
Vijay Konda and John Tsitsiklis · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A. McAllester, Satinder P. Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
J. Baxter and P. L. Bartlett · 2001
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Amir Beck and Marc Teboulle · 2003
Earlier work this paper cites.
The linear programming approach to approximate dynamic programming
D. P. De Farias and B. Van Roy · 2003
Earlier work this paper cites.
Error bounds for approximate policy iteration
Rémi Munos · 2003
Earlier work this paper cites.
Error bounds for approximate value iteration
Rémi Munos · 2005
Earlier work this paper cites.
An analysis of reinforcement learning with function approximation
Francisco S. Melo, Sean P. Meyn, and M. Isabel Ribeiro · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Natural actor-critic
Jan Peters and Stefan Schaal · 2008
Earlier work this paper cites.
Natural actor–critic algorithms
Shalabh Bhatnagar, Richard S. Sutton, Mohammad Ghavamzadeh, and Mark Lee · 2009
Earlier work this paper cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Richard Sutton, Hamid Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvari, and Eric Wiewiora · 2009
Earlier work this paper cites.
Dynamic Programming and Optimal Control: Volume II; Approximate Dynamic Programming
D. Bertsekas · 2012
Earlier work this paper cites.
Random design analysis of ridge regression
Daniel Hsu, Sham M. Kakade, and Tong Zhang · 2012
Earlier work this paper cites.
Non-strongly-convex smooth stochastic approximation with convergence rate o(1/n)
Francis Bach and Eric Moulines · 2013
Earlier work this paper cites.
Approximate policy iteration schemes: A comparison
Bruno Scherrer · 2014
Cited alongside, same era.
Understanding Machine Learning - From Theory to Algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
A stochastic quasi-newton method for large-scale optimization
R. H. Byrd, S. L. Hansen, Jorge Nocedal, and Y. Singer · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Cited alongside, same era.
Stochastic block BFGS: Squeezing more curvature out of data
Robert Gower, Donald Goldfarb, and Peter Richtarik · 2016
Cited alongside, same era.
Leverage the average: an analysis of kl regularization in reinforcement learning
Nino Vieillard, Tadashi Kozuno, Bruno Scherrer, Olivier Pietquin, Remi Munos, and Matthieu Geist · 2020
Later among the works it cites.
Neural policy gradient methods: Global optimality and rates of convergence
Lingxiao Wang, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2020
Later among the works it cites.
Improving sample complexity bounds for (natural) actor-critic algorithms
Tengyu Xu, Zhe Wang, and Yingbin Liang · 2020
Later among the works it cites.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
Lin Yang and Mengdi Wang · 2020
Later among the works it cites.
Variational policy gradient method for reinforcement learning with general utilities
Junyu Zhang, Alec Koppel, Amrit Singh Bedi, Csaba Szepesvari, and Mengdi Wang · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Analysis of classification-based policy iteration algorithms
Alessandro Lazaric, Mohammad Ghavamzadeh, and Rémi Munos · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
First-Order Methods in Optimization
Amir Beck · 2017
Cited alongside, same era.
Contextual decision processes with low Bellman rank are PAC-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E. Schapire · 2017
Cited alongside, same era.
A unified view of entropy-regularized markov decision processes, 2017
Gergely Neu, Anders Jonsson, and Vicenç Gómez · 2017
Cited alongside, same era.
Alekh Agarwal, Sham M. Kakade, Jason D. Lee, and Gaurav Mahajan · 2021
Later among the works it cites.
On the linear convergence of policy gradient methods for finite mdps
Jalaj Bhandari and Daniel Russo · 2021
Later among the works it cites.
Linear convergence of entropy-regularized natural policy gradient with linear function approximation, 2021
Semih Cayci, Niao He, and R. Srikant · 2021
Later among the works it cites.
Fast global convergence of natural policy gradient methods with entropy regularization
Shicong Cen, Chen Cheng, Yuxin Chen, Yuting Wei, and Yuejie Chi · 2021
Later among the works it cites.
On the linear convergence of natural policy gradient algorithm
Sajad Khodadadian, Prakirt Raj Jhunjhunwala, Sushil Mahavir Varma, and Siva Theja Maguluri · 2021
Later among the works it cites.
Leveraging non-uniformity in first-order non-convex optimization
Jincheng Mei, Yue Gao, Bo Dai, Csaba Szepesvari, and Dale Schuurmans · 2021
Later among the works it cites.
Cautiously optimistic policy optimization and exploration with linear function approximation
Andrea Zanette, Ching-An Cheng, and Alekh Agarwal · 2021
Later among the works it cites.
Policy mirror descent for regularized reinforcement learning: A generalized framework with linear convergence, 2021
Wenhao Zhan, Shicong Cen, Baihe Huang, Yuxin Chen, Jason D. Lee, and Yuejie Chi · 2021
Later among the works it cites.
On the convergence and sample efficiency of variance-reduced policy gradient method
Junyu Zhang, Chengzhuo Ni, Zheng Yu, Csaba Szepesvari, and Mengdi Wang · 2021
Later among the works it cites.
Linear convergence for natural policy gradient with log-linear policy parametrization, 2022
Carlo Alfano and Patrick Rebeschini · 2022
Closest in time.
Finite-time analysis of entropy-regularized neural natural actor-critic algorithm, 2022
Semih Cayci, Niao He, and R. Srikant · 2022
Closest in time.
Sample complexity of policy-based methods under off-policy sampling and linear function approximation
Zaiwei Chen and Siva Theja Maguluri · 2022
Closest in time.
Finite-sample analysis of off-policy natural actor–critic with linear function approximation
Zaiwei Chen, Sajad Khodadadian, and Siva Theja Maguluri · 2022
Closest in time.
On the global optimum convergence of momentum-based policy gradient
Yuhao Ding, Junzi Zhang, and Javad Lavaei · 2022
Closest in time.
Mirror learning: A unifying framework of policy optimisation
Jakub Grudzien, Christian A Schroeder De Witt, and Jakob Foerster · 2022
Closest in time.
Actor-critic is implicitly biased towards high entropy optimal policies
Yuzheng Hu, Ziwei Ji, and Matus Telgarsky · 2022
Closest in time.
Kl-entropy-regularized rl with a generative model is minimax optimal, 2022
Tadashi Kozuno, Wenhao Yang, Nino Vieillard, Toshinori Kitamura, Yunhao Tang, Jincheng Mei, Pierre Ménard, Mohammad Gheshlaghi Azar, Michal Valko, Rémi Munos, Olivier Pietquin, Matthieu Geist, and Csaba Szepesvári · 2022
Closest in time.
Policy mirror descent for reinforcement learning: linear convergence, new sampling complexity, and generalized problem classes
Guanghui Lan · 2022
Closest in time.
Stochastic linear optimization never overfits with quadratically-bounded losses on general data
Matus Telgarsky · 2022
Closest in time.
Mirror descent policy optimization
Manan Tomar, Lior Shani, Yonathan Efroni, and Mohammad Ghavamzadeh · 2022
Closest in time.
A general class of surrogate functions for stable and efficient reinforcement learning
Sharan Vaswani, Olivier Bachem, Simone Totaro, Robert Müller, Shivam Garg, Matthieu Geist, Marlos C. Machado, Pablo Samuel Castro, and Nicolas Le Roux · 2022
Closest in time.
On the convergence rates of policy gradient methods
Lin Xiao · 2022
Closest in time.
A general sample complexity analysis of vanilla policy gradient
Rui Yuan, Robert M. Gower, and Alessandro Lazaric · 2022
Closest in time.