Fetching the paper…
Reading the bibliography…
Explicit exploration in the action space was assumed to be indispensable for online policy gradient methods to avoid a drastic degradation in sample complexity, for solving general reinforcement learning problems over finite state and action spaces.
Problem complexity and method efficiency in optimization
Arkadij Semenovič Nemirovskij and David Borisovich Yudin · 1983
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai, Herbert Robbins, et al · 1985
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Reinforcement learning with replacing eligibility traces
Satinder P Singh and Richard S Sutton · 1996
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
John Tsitsiklis and Benjamin Van Roy · 1996
Earlier work this paper cites.
Actor-critic algorithms
Vijay Konda and John Tsitsiklis · 1999
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
On the convergence of optimistic policy iteration
John N Tsitsiklis · 2002
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
Peter Auer and Ronald Ortner · 2006
Earlier work this paper cites.
Pac model-free reinforcement learning
Alexander L Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L Littman · 2006
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Peter Auer, Thomas Jaksch, and Ronald Ortner · 2008
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
Optimization iii: Convex analysis, nonlinear programming theory, nonlinear programming algorithms
Aharon Ben-Tal and Arkadi Nemirovski · 2012
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Markov chains and mixing times
David A Levin and Yuval Peres · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
A finite time analysis of temporal difference learning with linear function approximation
Jalaj Bhandari, Daniel Russo, and Raghav Singal · 2018
Cited alongside, same era.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Provably correct optimization and exploration with non-linear policies
Fei Feng, Wotao Yin, Alekh Agarwal, and Lin Yang · 2021
Later among the works it cites.
Wenhao Zhan, Shicong Cen, Baihe Huang, Yuxin Chen, Jason D Lee, and Yuejie Chi · 2021
Later among the works it cites.
A natural actor-critic framework for zero-sum markov games
Ahmet Alacaoglu, Luca Viano, Niao He, and Volkan Cevher · 2022
Later among the works it cites.
Actor-critic is implicitly biased towards high entropy optimal policies
Yuzheng Hu, Ziwei Ji, and Matus Telgarsky · 2022
Later among the works it cites.
Finite sample analysis of two-time-scale natural actor-critic algorithm
Sajad Khodadadian, Thinh T Doan, Justin Romberg, and Siva Theja Maguluri · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Politex: Regret bounds for policy iteration using expert prediction
Yasin Abbasi-Yadkori, Peter Bartlett, Kush Bhatia, Nevena Lazic, Csaba Szepesvari, and Gellért Weisz · 2019
Cited alongside, same era.
Tsallis reinforcement learning: A unified framework for maximum entropy reinforcement learning
Kyungjae Lee, Sungyub Kim, Sungbin Lim, Sungjoon Choi, and Songhwai Oh · 2019
Cited alongside, same era.
Neural trust region/proximal policy optimization attains globally optimal policy
Boyi Liu, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2019
Cited alongside, same era.
Pc-pg: Policy cover directed exploration for provable policy gradient learning
Alekh Agarwal, Mikael Henaff, Sham Kakade, and Wen Sun · 2020
Cited alongside, same era.
Optimality and approximation with policy gradient methods in markov decision processes
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2020
Cited alongside, same era.
Provably efficient exploration in policy optimization
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2020
Cited alongside, same era.
Simple and optimal methods for stochastic variational inequalities, ii: Markovian noise and policy evaluation in reinforcement learning
Georgios Kotsalis, Guanghui Lan, and Tianjiao Li · 2022
Later among the works it cites.
Policy mirror descent for reinforcement learning: Linear convergence, new sampling complexity, and generalized problem classes
Guanghui Lan · 2022
Later among the works it cites.
Policy optimization over general state and action spaces
Guanghui Lan · 2022
Later among the works it cites.
Guanghui Lan, Yan Li, and Tuo Zhao · 2022
Later among the works it cites.
Stochastic first-order methods for average-reward markov decision processes
Tianjiao Li, Feiyang Wu, and Guanghui Lan · 2022
Later among the works it cites.
First-order policy optimization for robust markov decision process
Yan Li, Tuo Zhao, and Guanghui Lan · 2022
Later among the works it cites.
Yan Li, Tuo Zhao, and Guanghui Lan · 2022
Later among the works it cites.
Stochastic linear optimization never overfits with quadratically-bounded losses on general data
Matus Telgarsky · 2022
Later among the works it cites.
On the convergence rates of policy gradient methods
Lin Xiao · 2022
Later among the works it cites.