Fetching the paper…
Reading the bibliography…
We present new policy mirror descent (PMD) methods for solving reinforcement learning (RL) problems with either strongly convex or general convex regularizers.
Functional approximations and dynamic programming
R. Bellman and S. Dreyfus · 1959
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
A. S. Nemirovski and D. Yudin · 1983
Earlier work this paper cites.
A method for unconstrained convex minimization problem with the rate of convergence O ( 1 / k 2 ) O(1/k^{2})
Y. E. Nesterov · 1983
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R.S. Sutton, D. McAllester, S. Singh, and Y. Mansour · 1999
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
S. Kakade and J. Langford · 2002
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
A. Beck and M. Teboulle · 2003
Earlier work this paper cites.
Finite-Dimensional Variational Inequalities and Complementarity Problems, Volumes I and II
F. Facchinei and J. Pang · 2003
Earlier work this paper cites.
Online markov decision processes
Eyal Even-Dar, Sham. M. Kakade, and Yishay Mansour · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
A. S. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Cited alongside, same era.
On the convergence properties of non-euclidean extragradient methods for variational inequalities with generalized monotone operators
C. D. Dang and G. Lan · 2015
Cited alongside, same era.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
A. Agarwal, S. M. Kakade, J. D. Lee, and G. Mahajan · 2019
Cited alongside, same era.
Neural proximal/trust region policy optimization attains globally optimal policy
B. Liu, Q. Cai, Z. Yang, and Z. Wang · 2019
Cited alongside, same era.
A Note on the Linear Convergence of Policy Gradient Methods
Jalaj Bhandari and Daniel Russo · 2020
Cited alongside, same era.
First-order and Stochastic Optimization Methods for Machine Learning
G. Lan · 2020
Later among the works it cites.
On the Global Convergence Rates of Softmax Policy Gradient Methods
Jincheng Mei, Chenjun Xiao, Csaba Szepesvari, and Dale Schuurmans · 2020
Later among the works it cites.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps
Lior Shani, Yonathan Efroni, and Shie Mannor · 2020
Later among the works it cites.
Mirror descent policy optimization
M. Tomar, Lior Shani, Yonathan Efroni, and Mohammad Ghavamzadeh · 2020
Later among the works it cites.
Neural policy gradient methods: Global optimality and rates of convergence
L. Wang, Q. Cai, Zhuoran Yang, and Zhaoran Wang · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shicong Cen, Chen Cheng, Yuxin Chen, Yuting Wei, and Yuejie Chi · 2020
Cited alongside, same era.
Simple and optimal methods for stochastic variational inequalities, I: operator extrapolation
G. Kotsalis, G. Lan, and T. Li · 2020
Cited alongside, same era.
Simple and optimal methods for stochastic variational inequalities, II: Markovian noise and policy evaluation in reinforcement learning
G. Kotsalis, G. Lan, and T. Li · 2020
Cited alongside, same era.
Later among the works it cites.
Statistical estimation of ergodic markov chain kernel over discrete state space
G. Wolfer and A. Kontorovich · 2020
Later among the works it cites.
Improving sample complexity bounds for actor-critic algorithms
T. Xu, Zhe Wang, and Yingbin Liang · 2020
Later among the works it cites.
Finite-sample analysis of off-policy natural actor-critic algorithm
S. Khodadadian, Z. Chen, and S. T. Maguluri · 2021
Closest in time.