Fetching the paper…
Reading the bibliography…
Policy optimization methods are popular reinforcement learning algorithms, because their incremental and on-policy nature makes them more stable than the value-based counterparts.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Dimitri P Bertsekas, Dimitri P Bertsekas, Dimitri P Bertsekas, and Dimitri P Bertsekas · 1995
Earlier work this paper cites.
Temporal differences-based policy iteration and applications in neuro-dynamic programming
Dimitri P Bertsekas and Sergey Ioffe · 1996
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E Schapire · 1997
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, Yishay Mansour, et al · 1999
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup · 2000
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2001
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade et al · 2003
Earlier work this paper cites.
Provably efficient reinforcement learning for discounted mdps with feature mapping
Dongruo Zhou, Jiafan He, and Quanquan Gu · 2006
Earlier work this paper cites.
Pc-pg: Policy cover directed exploration for provable policy gradient learning
Alekh Agarwal, Mikael Henaff, Sham Kakade, and Wen Sun · 2007
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, T. Hayes, and Sham M. Kakade · 2008
Earlier work this paper cites.
Online markov decision processes
Eyal Even-Dar, Sham M Kakade, and Yishay Mansour · 2009
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2010
Earlier work this paper cites.
What are the statistical limits of offline rl with linear function approximation?
Ruosong Wang, Dean P Foster, and Sham M Kakade · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, David Pal, and Csaba Szepesvari · 2011
Earlier work this paper cites.
Dynamic policy programming with function approximation
Mohammad Gheshlaghi Azar, Bert Kappen, et al · 2011
Cited alongside, same era.
Contextual bandit algorithms with supervised learning guarantees
Alina Beygelzimer, John Langford, Lihong Li, Lev Reyzin, and Robert Schapire · 2011
Cited alongside, same era.
Zhuoran Yang, Chi Jin, Zhaoran Wang, Mengdi Wang, and Michael I Jordan · 2011
Cited alongside, same era.
Nearly minimax optimal reinforcement learning for linear mixture markov decision processes
Dongruo Zhou, Quanquan Gu, and Csaba Szepesvari · 2012
Cited alongside, same era.
Efficient exploration and value function generalization in deterministic systems
Zheng Wen and Benjamin Van Roy · 2013
Cited alongside, same era.
Global optimality guarantees for policy gradient methods
Jalaj Bhandari and Daniel Russo · 2019
Later among the works it cites.
Policy certificates: Towards accountable reinforcement learning
Christoph Dann, Lihong Li, Wei Wei, and Emma Brunskill · 2019
Later among the works it cites.
Is a good representation sufficient for sample efficient reinforcement learning?
Simon S Du, Sham M Kakade, Ruosong Wang, and Lin F Yang · 2019
Later among the works it cites.
Tight regret bounds for model-based reinforcement learning with greedy policies
Yonathan Efroni, Nadav Merlis, Mohammad Ghavamzadeh, and Shie Mannor · 2019
Later among the works it cites.
A theory of regularized markov decision processes
Matthieu Geist, Bruno Scherrer, and Olivier Pietquin · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Taming the monster: A fast and simple algorithm for contextual bandits
Alekh Agarwal, Daniel Hsu, Satyen Kale, John Langford, Lihong Li, and Robert Schapire · 2014
Cited alongside, same era.
Local policy search in a convex space and conservative policy iteration as boosted policy search
Bruno Scherrer and Matthieu Geist · 2014
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Pac reinforcement learning with rich observations
Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Remi Munos · 2017
Cited alongside, same era.
Contextual decision processes with low Bellman rank are PAC-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E. Schapire · 2017
Cited alongside, same era.
Fast rates with high probability in exp-concave statistical learning
Nishant Mehta · 2017
Cited alongside, same era.
Later among the works it cites.
Neural proximal/trust region policy optimization attains globally optimal policy
Boyi Liu, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2019
Later among the works it cites.
Worst-case regret bounds for exploration via randomized value functions
Daniel Russo · 2019
Later among the works it cites.
Optimism in reinforcement learning with generalized linear function approximation
Yining Wang, Ruosong Wang, Simon S Du, and Akshay Krishnamurthy · 2019
Later among the works it cites.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Andrea Zanette and Emma Brunskill · 2019
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvari, Mengdi Wang, and Lin F Yang · 2020
Later among the works it cites.
Provably efficient exploration in policy optimization
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2020
Later among the works it cites.
Agnostic q-learning with function approximation in deterministic systems: Tight bounds on approximation error and sample complexity, 2020
Simon S. Du, Jason D. Lee, Gaurav Mahajan, and Ruosong Wang · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Later among the works it cites.
Bandit Algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Later among the works it cites.
Optimistic policy optimization with bandit feedback
Lior Shani, Yonathan Efroni, Aviv Rosenberg, and Shie Mannor · 2020
Later among the works it cites.
Gellert Weisz, Philip Amortila, and Csaba Szepesvári · 2020
Later among the works it cites.
Reinforcement leaning in feature space: Matrix bandit, kernels, and regret bound
Lin F Yang and Mengdi Wang · 2020
Later among the works it cites.
Andrea Zanette · 2020
Later among the works it cites.
Zihan Zhang, Xiangyang Ji, and Simon S Du · 2020
Later among the works it cites.