Fetching the paper…
Reading the bibliography…
Despite the wide applications of Adam in reinforcement learning (RL), the theoretical convergence of Adam-type RL algorithms has not been established.
Finite-time analysis of Q-learning with linear function approximation
Chen, Z., Zhang, S., Doan, T. T., Maguluri, S. T., and Clarke, J.-P. (2019b) · 1905
Earlier work this paper cites.
Smoothing policies and safe policy gradients
Papini, M., Pirotta, M., and Restelli, M. (2019) · 1905
Earlier work this paper cites.
Global optimality guarantees for policy gradient methods
Bhandari, J. and Russo, D. (2019) · 1906
Earlier work this paper cites.
Global convergence of policy gradient methods to (almost) locally optimal policies
Zhang, K., Koppel, A., Zhu, H., and Başar, T. (2019) · 1906
Earlier work this paper cites.
Optimality and approximation with policy gradient methods in markov decision processes
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G. (2019) · 1908
Earlier work this paper cites.
Faster stochastic algorithms via history-gradient aided batch size adaptation
Ji, K., Wang, Z., Zhou, Y., and Liang, Y. (2019) · 1910
Earlier work this paper cites.
Natural actor-critic converges globally for hierarchical linear quadratic regulator
Luo, Y., Yang, Z., Wang, Z., and Kolar, M. (2019) · 1912
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S. (1988) · 1988
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P. (1992) · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
On-line Q-learning using connectionist systems
Rummery, G. A. and Niranjan, M. (1994) · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L. (1995) · 1995
Earlier work this paper cites.
On the averaged stochastic approximation for linear regression
Györfi, L. and Walk, H. (1996) · 1996
Earlier work this paper cites.
An analysis of temporal-diffference learning with function approximation
Tsitsiklis, J. N. and Van Roy, B. (1997) · 1997
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N. (2000) · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Baxter, J. and Bartlett, P. L. (2001) · 2001
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M. (2002) · 2002
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. (2002) · 2002
Earlier work this paper cites.
Sequential cost-sensitive decision making with reinforcement learning
Pednault, E., Abe, N., and Zadrozny, B. (2002) · 2002
Earlier work this paper cites.
daptive temporal difference learning with linear function approximation
Sun, T., Shen, H., Chen, T., and Li, D. (2020) · 2002
Cited alongside, same era.
Stochastic approximation and recursive algorithms and applications
Kushner, H. and Yin, G. G. (2003) · 2003
Cited alongside, same era.
Sensitivity and convergence of uniformly ergodic Markov chains
Mitrophanov, A. Y. (2005) · 2005
Cited alongside, same era.
Incremental natural actor-critic algorithms
Bhatnagar, S., Ghavamzadeh, M., Lee, M., and Sutton, R. S. (2008) · 2008
Cited alongside, same era.
Natural actor-critic algorithms
Bhatnagar, S., Sutton, R. S., Ghavamzadeh, M., and Lee, M. (2009) · 2009
Cited alongside, same era.
An actor-critic algorithm with function approximation for discounted cost constrained markov decision processes
Stochastic variance-reduced policy gradient
Papini, M., Binaghi, D., Canonaco, G., Pirotta, M., and Restelli, M. (2018) · 2018
Later among the works it cites.
On the convergence of Adam and beyond
Reddi, S. J., Kale, S., and Kumar, S. (2018) · 2018
Later among the works it cites.
Accelerated methods for deep reinforcement learning
Stooke, A. and Abbeel, P. (2018) · 2018
Later among the works it cites.
On the convergence of adaptive gradient methods for nonconvex optimization
Zhou, D., Tang, Y., Yang, Z., Cao, Y., and Gu, Q. (2018) · 2018
Later among the works it cites.
Neural temporal-difference learning converges to global optima
Cai, Q., Yang, Z., Lee, J. D., and Wang, Z. (2019) · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bhatnagar, S. (2010) · 2010
Cited alongside, same era.
Adaptive algorithms and stochastic approximations
Benveniste, A., Métivier, M., and Priouret, P. (2012) · 2012
Cited alongside, same era.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (2013) · 2013
Cited alongside, same era.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. (2014) · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2015) · 2015
Cited alongside, same era.
Policy gradient in lipschitz Markov decision processes
Pirotta, M., Restelli, M., and Bascetta, L. (2015) · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015) · 2015
Cited alongside, same era.
Castillo, G. A., Weng, B., Hereid, A., Wang, Z., and Zhang, W. (2019) · 2019
Later among the works it cites.
Characterizing the exact behaviors of temporal difference learning algorithms using Markov jump linear system theory
Hu, B. and Syed, U. (2019) · 2019
Later among the works it cites.
Non-asymptotic analysis of biased stochastic approximation scheme
Karimi, B., Miasojedow, B., Moulines, E., and Wai, H.-T. (2019) · 2019
Later among the works it cites.
Neural trust region/proximal policy optimization attains globally optimal policy
Liu, B., Cai, Q., Yang, Z., and Wang, Z. (2019) · 2019
Later among the works it cites.
Derivative-free methods for policy optimization: Guarantees for linear quadratic systems
Malik, D., Pananjady, A., Bhatia, K., Khamaru, K., Bartlett, P., and Wainwright, M. (2019) · 2019
Later among the works it cites.
Hessian aided policy gradient
Shen, Z., Ribeiro, A., Hassani, H., Qian, H., and Mi, C. (2019) · 2019
Later among the works it cites.
Finite-time error bounds for linear stochastic approximation andtd learning
Srikant, R. and Ying, L. (2019) · 2019
Later among the works it cites.
On the convergence proof of AMSGrad and a new version
Tran, P. T. and Phong, L. T. (2019) · 2019
Later among the works it cites.
The gap between model-based and model-free methods on the linear quadratic regulator: An asymptotic viewpoint
Tu, S. and Recht, B. (2019) · 2019
Later among the works it cites.
Provably global convergence of actor-critic: A case for linear quadratic regulator with ergodic cost
Yang, Z., Chen, Y., Hong, M., and Wang, Z. (2019) · 2019
Later among the works it cites.
Actor-critic provably finds nash equilibria of linear-quadratic mean-field games
Fu, Z., Yang, Z., Chen, Y., and Wang, Z. (2020) · 2020
Closest in time.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps
Shani, L., Efroni, Y., and Mannor, S. (2020) · 2020
Closest in time.
Neural policy gradient methods: Global optimality and rates of convergence
Wang, L., Cai, Q., Yang, Z., and Wang, Z. (2020) · 2020
Closest in time.
Sample efficient policy gradient methods with recursive variance reduction
Xu, P., Gao, F., and Gu, Q. (2020) · 2020
Closest in time.