Fetching the paper…
Reading the bibliography…
The softmax policy gradient (PG) method, which performs gradient ascent under softmax policy parameterization, is arguably one of the de facto implementations of policy optimization in modern reinforcement learning.
Global optimality guarantees for policy gradient methods
Bhandari, J. and Russo, D. (2019) · 1906
Earlier work this paper cites.
Temporal credit assignment in reinforcement learning
Sutton, R. S. (1984) · 1984
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N. (2000) · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M. (2002) · 2002
Earlier work this paper cites.
Non-asymptotic convergence analysis of two time-scale (natural) actor-critic algorithms
Xu, T., Wang, Z., and Liang, Y. (2020b) · 2005
Earlier work this paper cites.
Natural actor-critic
Peters, J. and Schaal, S. (2008) · 2008
Earlier work this paper cites.
Provable fictitious play for general mean-field games
Xie, Q., Yang, Z., Wang, Z., and Minca, A. (2020) · 2010
Earlier work this paper cites.
A primal approach to constrained policy optimization: Global optimality and finite-time analysis
Xu, T., Liang, Y., and Lan, G. (2020a) · 2011
Earlier work this paper cites.
Finding the near optimal policy via adaptive reduced regularization in MDPs
Yang, W., Li, X., Xie, G., and Zhang, Z. (2020) · 2011
Earlier work this paper cites.
Minimax PAC bounds on the sample complexity of reinforcement learning with a generative model
Azar, M. G., Munos, R., and Kappen, H. J. (2013) · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015) · 2015
Earlier work this paper cites.
Gradient descent only converges to minimizers
Lee, J. D., Simchowitz, M., Jordan, M. I., and Recht, B. (2016) · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016) · 2016
Earlier work this paper cites.
An alternative softmax operator for reinforcement learning
Asadi, K. and Littman, M. L. (2017) · 2017
Cited alongside, same era.
First-order methods in optimization
Beck, A. (2017) · 2017
Cited alongside, same era.
Gradient descent can take exponential time to escape saddle points
Du, S., Jin, C., Jordan, M., Póczos, B., Singh, A., and Lee, J. (2017) · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Cited alongside, same era.
Global convergence of policy gradient methods for the linear quadratic regulator
Fazel, M., Ge, R., Kakade, S., and Mesbahi, M. (2018) · 2018
Cited alongside, same era.
Neural proximal/trust region policy optimization attains globally optimal policy
A finite-time analysis of two time-scale actor-critic methods
Wu, Y. F., Zhang, W., Xu, P., and Gu, Q. (2020) · 2020
Later among the works it cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G. (2021) · 2021
Closest in time.
Global optimality of softmax policy gradient with single hidden layer neural networks in the mean-field regime
Agazzi, A. and Lu, J. (2021) · 2021
Closest in time.
On the linear convergence of policy gradient methods for finite MDPs
Bhandari, J. and Russo, D. (2021) · 2021
Closest in time.
Fast policy extragradient methods for competitive games with entropy regularization
Cen, S., Wei, Y., and Chi, Y. (2021) · 2021
Closest in time.
Episodic reinforcement learning in finite MDPs: Minimax lower bounds revisited
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liu, B., Cai, Q., Yang, Z., and Wang, Z. (2019) · 2019
Cited alongside, same era.
The gap between model-based and model-free methods on the linear quadratic regulator: An asymptotic viewpoint
Tu, S. and Recht, B. (2019) · 2019
Cited alongside, same era.
Neural policy gradient methods: Global optimality and rates of convergence
Wang, L., Cai, Q., Yang, Z., and Wang, Z. (2019) · 2019
Cited alongside, same era.
Optimization Foundations of Reinforcement Learning
Bhandari, J. (2020) · 2020
Cited alongside, same era.
Provably efficient exploration in policy optimization
Cai, Q., Yang, Z., Jin, C., and Wang, Z. (2020) · 2020
Cited alongside, same era.
Independent policy gradient methods for competitive reinforcement learning
Daskalakis, C., Foster, D. J., and Golowich, N. (2020) · 2020
Cited alongside, same era.
Natural policy gradient primal-dual method for constrained markov decision processes
Ding, D., Zhang, K., Basar, T., and Jovanovic, M. (2020) · 2020
Cited alongside, same era.
Domingues, O. D., Ménard, P., Kaufmann, E., and Valko, M. (2021) · 2021
Closest in time.
Is temporal difference learning optimal? an instance-dependent analysis
Khamaru, K., Pananjady, A., Ruan, F., Wainwright, M. J., and Jordan, M. I. (2021) · 2021
Closest in time.
Is Q-learning minimax optimal? a tight sample complexity analysis
Li, G., Cai, C., Chen, Y., Wei, Y., and Chi, Y. (2021) · 2021
Closest in time.
Leveraging non-uniformity in first-order non-convex optimization
Mei, J., Gao, Y., Dai, B., Szepesvari, C., and Schuurmans, D. (2021) · 2021
Closest in time.
Last-iterate convergence of decentralized optimistic gradient descent/ascent in infinite-horizon competitive Markov games
Wei, C.-Y., Lee, C.-W., Zhang, M., and Luo, H. (2021) · 2021
Closest in time.
Zhan, W., Cen, S., Huang, B., Chen, Y., Lee, J. D., and Chi, Y. (2021) · 2021
Closest in time.
Provably efficient policy gradient methods for two-player zero-sum Markov games
Zhao, Y., Tian, Y., Lee, J., and Du, S. (2021) · 2021
Closest in time.
A natural actor-critic framework for zero-sum Markov games
Alacaoglu, A., Viano, L., He, N., and Cevher, V. (2022) · 2022
Closest in time.
Finite sample analysis of two-time-scale natural actor-critic algorithm
Khodadadian, S., Doan, T. T., Romberg, J., and Maguluri, S. T. (2022) · 2022
Closest in time.
Model-based reinforcement learning is minimax-optimal for offline zero-sum Markov games
Yan, Y., Li, G., Chen, Y., and Fan, J. (2022) · 2022
Closest in time.