Fetching the paper…
Reading the bibliography…
Policy optimization is among the most popular and successful reinforcement learning algorithms, and there is increasing interest in understanding its theoretical guarantees.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2001
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Pac bounds for discounted MDPs
Tor Lattimore and Marcus Hutter · 2012
Earlier work this paper cites.
Stochastic shortest path problems under weak conditions
Dimitri P Bertsekas and Huizhen Yu · 2013
Earlier work this paper cites.
Online learning in episodic Markovian decision processes by relative entropy policy search
Alexander Zimin and Gergely Neu · 2013
Earlier work this paper cites.
Adaptivity and optimism: An improved exponentiated gradient algorithm
Jacob Steinhardt and Percy Liang · 2014
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Introduction to online convex optimization
Elad Hazan et al · 2016
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Earlier work this paper cites.
More adaptive algorithms for adversarial bandits
Chen-Yu Wei and Haipeng Luo · 2018
Cited alongside, same era.
Neural trust region/proximal policy optimization attains globally optimal policy
Boyi Liu, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2019
Cited alongside, same era.
Provably efficient exploration in policy optimization
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2020
Cited alongside, same era.
Near-optimal regret bounds for stochastic shortest path
Alon Cohen, Haim Kaplan, Yishay Mansour, and Aviv Rosenberg · 2020
Cited alongside, same era.
Dynamic regret of policy optimization in non-stationary environments
Yingjie Fei, Zhuoran Yang, Zhaoran Wang, and Qiaomin Xie · 2020
Cited alongside, same era.
Optimistic policy optimization with bandit feedback
Lior Shani, Yonathan Efroni, Aviv Rosenberg, and Shie Mannor · 2020
Cited alongside, same era.
Minimax regret for stochastic shortest path
Alon Cohen, Yonathan Efroni, Yishay Mansour, and Aviv Rosenberg · 2021
Later among the works it cites.
Online learning for stochastic shortest path model via posterior sampling
Mehdi Jafarnia-Jahromi, Liyu Chen, Rahul Jain, and Haipeng Luo · 2021
Later among the works it cites.
Improved regret analysis for variance-adaptive linear bandits and horizon-free linear mixture MDPs
Yeoneung Kim, Insoon Yang, and Kwang-Sung Jun · 2021
Later among the works it cites.
Policy optimization in adversarial MDPs: Improved exploration via dilated bonuses
Haipeng Luo, Chen-Yu Wei, and Chung-Wei Lee · 2021
Later among the works it cites.
Learning stochastic shortest path with linear function approximation
Yifei Min, Jiafan He, Tianhao Wang, and Quanquan Gu · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
No-regret exploration in goal-oriented reinforcement learning
Jean Tarbouriech, Evrard Garcelon, Michal Valko, Matteo Pirotta, and Alessandro Lazaric · 2020
Cited alongside, same era.
Neural policy gradient methods: Global optimality and rates of convergence
Lingxiao Wang, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2020
Cited alongside, same era.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2021
Cited alongside, same era.
Finding the stochastic shortest path with low regret: The adversarial cost and unknown transition case
Liyu Chen and Haipeng Luo · 2021
Cited alongside, same era.
Implicit finite-horizon approximation and efficient optimal algorithms for stochastic shortest path
Liyu Chen, Mehdi Jafarnia-Jahromi, Rahul Jain, and Haipeng Luo
Cited in the paper.
Improved no-regret algorithms for stochastic shortest path with linear MDP
Liyu Chen, Rahul Jain, and Haipeng Luo
Cited in the paper.
Later among the works it cites.
Stochastic shortest path with adversarially changing costs
Aviv Rosenberg and Yishay Mansour · 2021
Later among the works it cites.
Stochastic shortest path: Minimax, parameter-free and towards horizon-free regret
Jean Tarbouriech, Runlong Zhou, Simon S Du, Matteo Pirotta, Michal Valko, and Alessandro Lazaric · 2021
Later among the works it cites.
Nearly optimal policy optimization with stable at any time guarantee
Tianhao Wu, Yunchang Yang, Han Zhong, Liwei Wang, Simon S Du, and Jiantao Jiao · 2021
Later among the works it cites.
Variance-aware confidence set: Variance-dependent bound for linear bandits and horizon-free bound for linear mixture MDP
Zihan Zhang, Jiaqi Yang, Xiangyang Ji, and Simon S Du · 2021
Later among the works it cites.
Nearly minimax optimal reinforcement learning for linear mixture Markov decision processes
Dongruo Zhou, Quanquan Gu, and Csaba Szepesvari · 2021
Later among the works it cites.