Fetching the paper…
Reading the bibliography…
Policy gradient (PG) methods are popular and efficient for large-scale reinforcement learning due to their relative stability and incremental nature.
On the convergence of adam and beyond
Reddi, S. J., Kale, S., and Kumar, S. (2019) · 1904
Earlier work this paper cites.
Momentum-based variance reduction in non-convex sgd
Cutkosky, A. and Orabona, F. (2019) · 1905
Earlier work this paper cites.
Global optimality guarantees for policy gradient methods
Bhandari, J. and Russo, D. (2019) · 1906
Earlier work this paper cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G. (2019) · 1908
Earlier work this paper cites.
Neural policy gradient methods: Global optimality and rates of convergence
Wang, L., Cai, Q., Yang, Z., and Wang, Z. (2019) · 1909
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Constrained Markov decision processes
Altman, E. (1999) · 1999
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Qian, N. (1999) · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., Mansour, Y., et al. (1999) · 1999
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N. (2000) · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Baxter, J. and Bartlett, P. L. (2001) · 2001
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M. (2001) · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J. (2002) · 2002
Earlier work this paper cites.
Non-asymptotic convergence of adam-type reinforcement learning algorithms under markovian sampling
Xiong, H., Xu, T., Liang, Y., and Zhang, W. (2020) · 2002
Earlier work this paper cites.
Exploration-exploitation in constrained mdps
Efroni, Y., Mannor, S., and Pirotta, M. (2020) · 2003
Earlier work this paper cites.
On actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N. (2003) · 2003
Earlier work this paper cites.
Stochastic recursive momentum for policy gradient methods
Yuan, H., Lian, X., Liu, J., and Zhou, Y. (2020) · 2003
Earlier work this paper cites.
Non-asymptotic convergence analysis of two time-scale (natural) actor-critic algorithms
Xu, T., Wang, Z., and Liang, Y. (2020d) · 2005
Earlier work this paper cites.
Hong, M., Wai, H.-T., Wang, Z., and Yang, Z. (2020) · 2007
Cited alongside, same era.
Single-timescale actor-critic provably finds globally optimal policy
Fu, Z., Yang, Z., and Wang, Z. (2020) · 2008
Cited alongside, same era.
Natural actor-critic
Peters, J. and Schaal, S. (2008) · 2008
Cited alongside, same era.
Natural actor–critic algorithms
Bhatnagar, S., Sutton, R. S., Ghavamzadeh, M., and Lee, M. (2009) · 2009
Cited alongside, same era.
Learning bounds for importance weighting
Cortes, C., Mansour, Y., and Mohri, M. (2010) · 2010
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Global convergence of policy gradient methods for the linear quadratic regulator
Fazel, M., Ge, R., Kakade, S., and Mesbahi, M. (2018) · 2018
Later among the works it cites.
Stochastic variance-reduced policy gradient
Papini, M., Binaghi, D., Canonaco, G., Pirotta, M., and Restelli, M. (2018) · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
Variance reduction for policy gradient with action-dependent factorized baselines
Wu, C., Rajeswaran, A., Duan, Y., Kumar, V., Bayen, A. M., Kakade, S., Mordatch, I., and Abbeel, P. (2018) · 2018
Later among the works it cites.
Neural-symbolic vqa: Disentangling reasoning from vision and language understanding
Yi, K., Wu, J., Gan, C., Torralba, A., Kohli, P., and Tenenbaum, J. B. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Johnson, R. and Zhang, T. (2013) · 2013
Cited alongside, same era.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. (2014) · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2015) · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2015) · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015) · 2015
Cited alongside, same era.
Variance reduction for faster non-convex optimization
Allen-Zhu, Z. and Hazan, E. (2016) · 2016
Cited alongside, same era.
Stochastic variance reduction for nonconvex optimization
Reddi, S. J., Hefny, A., Sra, S., Poczos, B., and Smola, A. (2016) · 2016
Cited alongside, same era.
Neural trust region/proximal policy optimization attains globally optimal policy
Liu, B., Cai, Q., Yang, Z., and Wang, Z. (2019) · 2019
Later among the works it cites.
Hessian aided policy gradient
Shen, Z., Ribeiro, A., Hassani, H., Qian, H., and Mi, C. (2019) · 2019
Later among the works it cites.
Natural policy gradient primal-dual method for constrained markov decision processes
Ding, D., Zhang, K., Basar, T., and Jovanovic, M. R. (2020) · 2020
Later among the works it cites.
Momentum-based policy gradient methods
Huang, F., Gao, S., Pei, J., and Huang, H. (2020) · 2020
Later among the works it cites.
Graph policy gradients for large scale robot control
Khan, A., Tolstaya, E., Ribeiro, A., and Kumar, V. (2020) · 2020
Later among the works it cites.
An improved analysis of (variance-reduced) policy gradient and natural policy gradient methods
Liu, Y., Zhang, K., Basar, T., and Yin, W. (2020) · 2020
Later among the works it cites.
On the global convergence rates of softmax policy gradient methods
Mei, J., Xiao, C., Szepesvari, C., and Schuurmans, D. (2020) · 2020
Later among the works it cites.
A hybrid stochastic policy gradient algorithm for reinforcement learning
Pham, N., Nguyen, L., Phan, D., Nguyen, P. H., Dijk, M., and Tran-Dinh, Q. (2020) · 2020
Later among the works it cites.
Cooperative multiagent deep deterministic policy gradient (comaddpg) for intelligent connected transportation with unsignalized intersection
Wu, T., Jiang, M., and Zhang, L. (2020a) · 2020
Later among the works it cites.
Leveraging non-uniformity in first-order non-convex optimization
Mei, J., Gao, Y., Dai, B., Szepesvari, C., and Schuurmans, D. (2021) · 2021
Closest in time.
A dual approach to constrained markov decision processes with entropy regularization
Ying, D., Ding, Y., and Lavaei, J. (2021) · 2021
Closest in time.
A general sample complexity analysis of vanilla policy gradient
Yuan, R., Gower, R. M., and Lazaric, A. (2021) · 2021
Closest in time.
Cautiously optimistic policy optimization and exploration with linear function approximation
Zanette, A., Cheng, C.-A., and Agarwal, A. (2021) · 2021
Closest in time.