Fetching the paper…
Reading the bibliography…
This paper revisits the temporal difference (TD) learning algorithm for the policy evaluation tasks in reinforcement learning.
R. Mendelssohn, “An iterative aggregation procedure for markov decision processes,” Operations Research
1982
Earlier work this paper cites.
R. S. Sutton, “Learning to predict by the methods of temporal differences,” Machine learning
1988
Earlier work this paper cites.
D. P. Bertsekas and D. A. Castanon, “Adaptive aggregation methods for infinite horizon dynamic programming,” IEEE Transactions on Automatic Control
1989
Earlier work this paper cites.
T. Jaakkola, M. I. Jordan, and S. P. Singh, “Convergence of stochastic iterative dynamic programming algorithms,” in Advances in neural information processing systems
1994
Earlier work this paper cites.
L. Baird, “Residual algorithms: Reinforcement learning with function approximation,” in Machine Learning
1995
Earlier work this paper cites.
R. S. Sutton, “Generalization in reinforcement learning: Successful examples using sparse coarse coding,” Advances in neural information processing systems
1996
Earlier work this paper cites.
J. N. Tsitsiklis and B. Van Roy, “An analysis of temporal-difference learning with function approximation,” IEEE transactions on automatic control
1997
Earlier work this paper cites.
MIT press Cambridge, 1998
R. S. Sutton, A. G. Barto, et al · 1998
Earlier work this paper cites.
PhD thesis, Massachusetts Institute of Technology, 1998
B. Van Roy, Learning and value function approximation in complex decision processes · 1998
Earlier work this paper cites.
V. S. Borkar and S. P. Meyn, “The ode method for convergence of stochastic approximation and reinforcement learning,” SIAM Journal on Control and Optimization
2000
Earlier work this paper cites.
E. Even-Dar and Y. Mansour, “Learning rates for Q-learning,” Journal of machine learning Research
2003
Earlier work this paper cites.
L. Li, T. J. Walsh, and M. L. Littman, “Towards a unified theory of state abstraction for mdps.,” ISAIM
2006
Earlier work this paper cites.
B. Van Roy, “Performance loss bounds for approximate value iteration with state aggregation,” Mathematics of Operations Research
2006
Earlier work this paper cites.
R. S. Sutton, H. R. Maei, and C. Szepesvári, “A convergent o ( n ) o(n) temporal-difference algorithm for off-policy learning with linear function approximation,” in Advances in Neural Information Processing Systems
2009
Earlier work this paper cites.
H. B. McMahan and M. Streeter, “Adaptive bound optimization for online convex optimization,” COLT 2010
2010
Earlier work this paper cites.
J. Duchi, E. Hazan, and Y. Singer, “Adaptive subgradient methods for online learning and stochastic optimization,” Journal of Machine Learning Research
2011
Earlier work this paper cites.
D. P. Bertsekas, “Dynamic programming and optimal control 3rd edition, volume ii,” Belmont, MA: Athena Scientific
2011
Cited alongside, same era.
T. Tieleman and G. Hinton, “Rmsprop: Neural networks for machine learning,” University of Toronto, Technical Report
2012
Cited alongside, same era.
M. D. Zeiler, “ADADELTA: An adaptive learning rate method,” arXiv preprint:1212.5701
2012
Cited alongside, same era.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” ICLR
2014
Cited alongside, same era.
B. Liu, J. Liu, M. Ghavamzadeh, S. Mahadevan, and M. Petrik, “Finite-sample analysis of proximal gradient td algorithms.,” in Proc. Conf. Uncertainty in Artificial Intelligence
2015
Cited alongside, same era.
X. Chen, S. Liu, R. Sun, and M. Hong, “On the convergence of a class of adam-type algorithms for non-convex optimization,” ICLR
2018
Later among the works it cites.
Z. Chen, Z. Yuan, J. Yi, B. Zhou, E. Chen, and T. Yang, “Universal stagewise learning for non-convex problems with convergence on averaged solutions,” in International Conference on Learning Representations
2018
Later among the works it cites.
R. Srikant and L. Ying, “Finite-time error bounds for linear stochastic approximation and TD learning,” COLT
2019
Later among the works it cites.
B. Hu and U. Syed, “Characterizing the exact behaviors of temporal difference learning algorithms using markov jump linear system theory,” in Advances in Neural Information Processing Systems
2019
Later among the works it cites.
T. Doan, S. Maguluri, and J. Romberg, “Finite-time analysis of distributed TD(0) with linear function approximation on multi-agent reinforcement learning,” in International Conference on Machine Learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al
2015
Cited alongside, same era.
T. Dozat, “Incorporating nesterov momentum into adam,” 2016
2016
Cited alongside, same era.
D. Abel, D. Hershkowitz, and M. Littman, “Near optimal behavior via approximate state abstraction,” in International Conference on Machine Learning
2016
Cited alongside, same era.
2016
Cited alongside, same era.
A. Devraj and S. Meyn, “Zap Q-learning,” in Advances in Neural Information Processing Systems
2017
Cited alongside, same era.
American Mathematical Soc., 2017
D. A. Levin and Y. Peres, Markov chains and mixing times · 2017
Cited alongside, same era.
R. Lowe, Y. I. Wu, A. Tamar, J. Harb, O. P. Abbeel, and I. Mordatch, “Multi-agent actor-critic for mixed cooperative-competitive environments,” in Advances in neural information processing systems
2017
Cited alongside, same era.
2019
Later among the works it cites.
L. Y. Harsh Gupta, R. Srikant, “Finite-time performance bounds and adaptive learning rate selection for two time-scale reinforcement learning,” in Advances in Neural Information Processing Systems
2019
Later among the works it cites.
X. Li and F. Orabona, “On the convergence of stochastic gradient descent with adaptive stepsizes,” in The 22nd International Conference on Artificial Intelligence and Statistics
2019
Later among the works it cites.
R. Ward, X. Wu, and L. Bottou, “Adagrad stepsizes: Sharp convergence over nonconvex landscapes,” in International Conference on Machine Learning
2019
Later among the works it cites.
S. J. Reddi, S. Kale, and S. Kumar, “On the convergence of adam and beyond,” ICLR
2019
Later among the works it cites.
F. Zou, L. Shen, Z. Jie, W. Zhang, and W. Liu, “A sufficient condition for convergences of adam and rmsprop,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
2019
Later among the works it cites.
M. Liu, Y. Mroueh, J. Ross, W. Zhang, X. Cui, P. Das, and T. Yang, “Towards better understanding of adaptive gradient algorithms in generative adversarial nets,” in International Conference on Learning Representations
2019
Later among the works it cites.
Y. Duan, T. Ke, and M. Wang, “State aggregation learning from markov transition data,” Advances in Neural Information Processing Systems
2019
Later among the works it cites.
N. Vieillard, B. Scherrer, O. Pietquin, and M. Geist, “Momentum in reinforcement learning,” in International Conference on Artificial Intelligence and Statistics
2020
Closest in time.
L. Shani, Y. Efroni, and S. Mannor, “Adaptive trust region policy optimization: Global convergence and faster rates for regularized MDPs,” in Proceedings of the AAAI Conference on Artificial Intelligence
2020
Closest in time.