Fetching the paper…
Reading the bibliography…
Policy optimization, which finds the desired policy by maximizing value functions via optimization techniques, lies at the heart of reinforcement learning (RL).
Tsallis reinforcement learning: A unified framework for maximum entropy reinforcement learning
Lee, K., Kim, S., Lim, S., Choi, S., and Oh, S. (2019) · 1902
Earlier work this paper cites.
Global optimality guarantees for policy gradient methods
Bhandari, J. and Russo, D. (2019) · 1906
Earlier work this paper cites.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized MDPs
Shani, L., Efroni, Y., and Mannor, S. (2019) · 1909
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Nemirovsky, A. S. and Yudin, D. B. (1983) · 1983
Earlier work this paper cites.
Possible generalization of Boltzmann-Gibbs statistics
Tsallis, C. (1988) · 1988
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Williams, R. J. and Peng, J. (1991) · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Proximal minimization methods with generalized Bregman functions
Kiwiel, K. C. (1997) · 1997
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M. (2002) · 2002
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Beck, A. and Teboulle, M. (2003) · 2003
Earlier work this paper cites.
Exploration-exploitation in constrained MDPs
Efroni, Y., Mannor, S., and Pirotta, M. (2020) · 2003
Earlier work this paper cites.
Mirror descent policy optimization
Tomar, M., Shani, L., Efroni, Y., and Ghavamzadeh, M. (2020) · 2005
Earlier work this paper cites.
A note on the linear convergence of policy gradient methods
Bhandari, J. and Russo, D. (2020) · 2007
Earlier work this paper cites.
Agazzi, A. and Lu, J. (2020) · 2010
Earlier work this paper cites.
Composite objective mirror descent
Duchi, J. C., Shalev-Shwartz, S., Singer, Y., and Tewari, A. (2010) · 2010
Earlier work this paper cites.
Sample efficient reinforcement learning with REINFORCE
Zhang, J., Kim, J., O’Donoghue, B., and Boyd, S. (2020a) · 2010
Earlier work this paper cites.
Primal-dual first-order methods with O ( 1 / ε ) {O(1/\varepsilon)} iteration-complexity for cone programming
Lan, G., Lu, Z., and Monteiro, R. D. (2011) · 2011
Earlier work this paper cites.
A primal approach to constrained policy optimization: Global optimality and finite-time analysis
Xu, T., Liang, Y., and Lan, G. (2020) · 2011
Cited alongside, same era.
Safe exploration in markov decision processes
Moldovan, T. M. and Abbeel, P. (2012) · 2012
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L. (2014) · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015) · 2015
Cited alongside, same era.
Concrete problems in ai safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D. (2016) · 2016
Neural trust region/proximal policy optimization attains globally optimal policy
Liu, B., Cai, Q., Yang, Z., and Wang, Z. (2019) · 2019
Later among the works it cites.
Neural policy gradient methods: Global optimality and rates of convergence
Wang, L., Cai, Q., Yang, Z., and Wang, Z. (2019) · 2019
Later among the works it cites.
Sample efficient policy gradient methods with recursive variance reduction
Xu, P., Gao, F., and Gu, Q. (2019) · 2019
Later among the works it cites.
Convergent policy optimization for safe reinforcement learning
Yu, M., Yang, Z., Kolar, M., and Wang, Z. (2019) · 2019
Later among the works it cites.
Optimality and approximation with policy gradient methods in Markov decision processes
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G. (2020) · 2020
Later among the works it cites.
An improved analysis of (variance-reduced) policy gradient and natural policy gradient methods
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
First-order methods in optimization
Beck, A. (2017) · 2017
Cited alongside, same era.
Dynamic programming and optimal control (4th edition)
Bertsekas, D. P. (2017) · 2017
Cited alongside, same era.
A unified view of entropy-regularized Markov decision processes
Neu, G., Jonsson, A., and Gómez, V. (2017) · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Cited alongside, same era.
SBEED: Convergent reinforcement learning with nonlinear function approximation
Dai, B., Shaw, A., Li, L., Xiao, L., He, N., Liu, Z., Chen, J., and Song, L. (2018) · 2018
Cited alongside, same era.
Global convergence of policy gradient methods for the linear quadratic regulator
Fazel, M., Ge, R., Kakade, S., and Mesbahi, M. (2018) · 2018
Cited alongside, same era.
Liu, Y., Zhang, K., Basar, T., and Yin, W. (2020) · 2020
Later among the works it cites.
Leverage the average: an analysis of KL regularization in reinforcement learning
Vieillard, N., Kozuno, T., Scherrer, B., Pietquin, O., Munos, R., and Geist, M. (2020) · 2020
Later among the works it cites.
Fast policy extragradient methods for competitive games with entropy regularization
Cen, S., Wei, Y., and Chi, Y. (2021) · 2021
Closest in time.
Provably efficient safe exploration via primal-dual policy optimization
Ding, D., Wei, X., Yang, Z., Wang, Z., and Jovanovic, M. (2021) · 2021
Closest in time.
Adaptive approximate policy iteration
Hao, B., Lazic, N., Abbasi-Yadkori, Y., Joulani, P., and Szepesvári, C. (2021) · 2021
Closest in time.
On the linear convergence of natural policy gradient algorithm
Khodadadian, S., Jhunjhunwala, P. R., Varma, S. M., and Maguluri, S. T. (2021) · 2021
Closest in time.
Improved regret bound and experience replay in regularized policy iteration
Lazic, N., Yin, D., Abbasi-Yadkori, Y., and Szepesvari, C. (2021) · 2021
Closest in time.
Leveraging non-uniformity in first-order non-convex optimization
Mei, J., Gao, Y., Dai, B., Szepesvari, C., and Schuurmans, D. (2021) · 2021
Closest in time.
Global convergence of policy gradient for linear-quadratic mean-field control/game in continuous time
Wang, W., Han, J., Yang, Z., and Wang, Z. (2021) · 2021
Closest in time.
Policy mirror descent for reinforcement learning: Linear convergence, new sampling complexity, and generalized problem classes
Lan, G. (2022) · 2022
Closest in time.
On the convergence rates of policy gradient methods
Xiao, L. (2022) · 2022
Closest in time.
Provably efficient policy optimization for two-player zero-sum markov games
Zhao, Y., Tian, Y., Lee, J., and Du, S. (2022) · 2022
Closest in time.
Softmax policy gradient methods can take exponential time to converge
Li, G., Wei, Y., Chi, Y., and Chen, Y. (2023) · 2023
Closest in time.