Fetching the paper…
Reading the bibliography…
We analyze the convergence rate of the unregularized natural policy gradient algorithm with log-linear policy parametrizations in infinite-horizon discounted Markov decision processes.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David A. McAllester, Satinder P. Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langfor · 2002
Earlier work this paper cites.
A natural policy gradient
Sham M. Kakade · 2002
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Sébastien Bubeck · 2015
Earlier work this paper cites.
The information geometry of mirror descent
Garvesh Raskutti and Sayan Mukherjee · 2015
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
A finite time analysis of temporal difference learning with linear function approximation
Jalaj Bhandari, Daniel Russo, and Raghav Singal · 2018
Earlier work this paper cites.
Flambe: Structural complexity and representation learning of low rank mdps
Alekh Agarwal, Sham Kakade, Akshay Krishnamurthy, and Wen Sun · 2020
Earlier work this paper cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I. Jordan · 2020
Earlier work this paper cites.
On the global convergence rates of softmax policy gradient methods
Jincheng Mei, Chenjun Xiao, Csaba Szepesvari, and Dale Schuurmans · 2020
Earlier work this paper cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2021
Cited alongside, same era.
On the linear convergence of policy gradient methods for finite mdps, 2021
Jalaj Bhandari and Daniel Russo · 2021
Cited alongside, same era.
Linear convergence of entropy-regularized natural policy gradient with linear function approximation
Semih Cayci, Niao He, and R Srikant · 2021
Cited alongside, same era.
Fast global convergence of natural policy gradient methods with entropy regularization
Shicong Cen, Chen Cheng, Yuxin Chen, Yuting Wei, and Yuejie Chi · 2021
Cited alongside, same era.
Actor-critic is implicitly biased towards high entropy optimal policies
Yuzheng Hu, Ziwei Ji, and Matus Telgarsky · 2021
Cited alongside, same era.
Representation learning for online and offline rl in low-rank mdps
Masatoshi Uehara, Xuezhou Zhang, and Wen Sun · 2021
Later among the works it cites.
Cautiously optimistic policy optimization and exploration with linear function approximation
Andrea Zanette, Ching-An Cheng, and Alekh Agarwal · 2021
Later among the works it cites.
Wenhao Zhan, Shicong Cen, Baihe Huang, Yuxin Chen, Jason D Lee, and Yuejie Chi · 2021
Later among the works it cites.
Policy mirror descent for reinforcement learning: Linear convergence, new sampling complexity, and generalized problem classes
Guanghui Lan · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the linear convergence of natural policy gradient algorithm
Sajad Khodadadian, Prakirt Raj Jhunjhunwala, Sushil Mahavir Varma, and Siva Theja Maguluri · 2021
Cited alongside, same era.
Gen Li, Yuxin Chen, Yuejie Chi, Yuantao Gu, and Yuting Wei · 2021
Cited alongside, same era.
Model-free representation learning and exploration in low-rank mdps
Aditya Modi, Jinglin Chen, Akshay Krishnamurthy, Nan Jiang, and Alekh Agarwal · 2021
Cited alongside, same era.
On finite-time convergence of actor-critic algorithm
Shuang Qiu, Zhuoran Yang, Jieping Ye, and Zhaoran Wang · 2021
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz
Cited in the paper.
Yan Li, Tuo Zhao, and Guanghui Lan · 2022
Closest in time.
First-order regret in reinforcement learning with linear function approximation: A robust estimation approach
Andrew J Wagenmaker, Yifang Chen, Max Simchowitz, Simon Du, and Kevin Jamieson · 2022
Closest in time.
On the convergence rates of policy gradient methods
Lin Xiao · 2022
Closest in time.
Making linear mdps practical via contrastive representation learning, 2022
Tianjun Zhang, Tongzheng Ren, Mengjiao Yang, Joseph Gonzalez, Dale Schuurmans, and Bo Dai · 2022
Closest in time.
Linear convergence of natural policy gradient methods with log-linear policies
Rui Yuan, Simon Shaolei Du, Robert M. Gower, Alessandro Lazaric, and Lin Xiao · 2023
Closest in time.