Fetching the paper…
Reading the bibliography…
We revisit the finite time analysis of policy gradient methods in the one of the simplest settings: finite state and action MDPs with a policy class consisting of all stochastic policies and with exact gradient evaluations.
Efficient exploration in reinforcement learning
Sebastian B Thrun · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Dynamic programming and optimal control
Dimitri P Bertsekas · 1995
Earlier work this paper cites.
Direct gradient-based reinforcement learning: I. gradient estimation algorithms
Jonathan Baxter and Peter L Bartlett · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Simulation-based optimization of markov reward processes
Peter Marbach and John N Tsitsiklis · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2002
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Amir Beck and Marc Teboulle · 2003
Earlier work this paper cites.
Efficient projections onto the l 1-ball for learning in high dimensions
John Duchi, Shai Shalev-Shwartz, Yoram Singer, and Tushar Chandra · 2008
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Cited alongside, same era.
Local policy search in a convex space and conservative policy iteration as boosted policy search
Bruno Scherrer and Matthieu Geist · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
Convex optimization: Algorithms and complexity
Sébastien Bubeck et al · 2015
Cited alongside, same era.
The information geometry of mirror descent
Garvesh Raskutti and Sayan Mukherjee · 2015
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Politex: Regret bounds for policy iteration using expert prediction
Yasin Abbasi-Yadkori, Peter Bartlett, Kush Bhatia, Nevena Lazic, Csaba Szepesvari, and Gellért Weisz · 2019
Later among the works it cites.
Global optimality guarantees for policy gradient methods
Jalaj Bhandari and Daniel Russo · 2019
Later among the works it cites.
A theory of regularized Markov decision processes
Matthieu Geist, Bruno Scherrer, and Olivier Pietquin · 2019
Later among the works it cites.
Optimality and approximation with policy gradient methods in markov decision processes
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2020
Closest in time.
Fast global convergence of natural policy gradient methods with entropy regularization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
A unified view of entropy-regularized markov decision processes
Gergely Neu, Anders Jonsson, and Vicenç Gómez · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Shicong Cen, Chen Cheng, Yuxin Chen, Yuting Wei, and Yuejie Chi · 2020
Closest in time.
Provably efficient adaptive approximate policy iteration
Botao Hao, Nevena Lazic, Yasin Abbasi-Yadkori, Pooria Joulani, and Csaba Szepesvari · 2020
Closest in time.
On the global convergence rates of softmax policy gradient methods
Jincheng Mei, Chenjun Xiao, Csaba Szepesvari, and Dale Schuurmans · 2020
Closest in time.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps
Lior Shani, Yonathan Efroni, and Shie Mannor · 2020
Closest in time.