Fetching the paper…
Reading the bibliography…
Existing studies indicate that momentum ideas in conventional optimization can be used to improve the performance of Q-learning algorithms.
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
Q-learning
Christopher J.C.H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Convergence of stochastic iterative dynamic programming algorithms
Tommi Jaakkola, Michael I Jordan, and Satinder P Singh · 1994
Earlier work this paper cites.
Asynchronous stochastic approximation and Q-learning
John N Tsitsiklis · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
Neuro-Dynamic Programming , volume 5
Dimitri P. Bertsekas and John N Tsitsiklis · 1996
Earlier work this paper cites.
The asymptotic convergence-rate of Q-learning
Csaba Szepesvári · 1998
Earlier work this paper cites.
Finite-sample convergence rates for Q-learning and indirect algorithms
Michael J Kearns and Satinder P Singh · 1999
Earlier work this paper cites.
The ode method for convergence of stochastic approximation and reinforcement learning
Vivek S Borkar and Sean P Meyn · 2000
Earlier work this paper cites.
Convergence of Q-learning: A simple proof
Francisco S Melo · 2001
Earlier work this paper cites.
Learning rates for Q-learning
Eyal Even-Dar and Yishay Mansour · 2003
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications , volume 35
Harold Kushner and G George Yin · 2003
Earlier work this paper cites.
Q-learning with linear function approximation
Francisco S Melo and M Isabel Ribeiro · 2007
Earlier work this paper cites.
The probabilistic method
Noga Alon and Joel H. Spencer · 2008
Cited alongside, same era.
Speedy Q-learning
Mohammad Gheshlaghi Azar, Remi Munos, M Ghavamzadaeh, and Hilbert J Kappen · 2011
Cited alongside, same era.
Playing Atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Introductory lectures on convex optimization: A basic course , volume 87
Yurii Nesterov · 2013
Cited alongside, same era.
Zap Q-learning
Adithya M Devraj and Sean Meyn · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis · 2017
Provably efficient Q-learning with function approximation via distribution shift error checking oracle
Simon S Du, Yuping Luo, Ruosong Wang, and Hanrui Zhang · 2019
Later among the works it cites.
A theoretical analysis of deep Q-learning
Jianqing Fan, Zhaoran Wang, Yuchen Xie, and Zhuoran Yang · 2019
Later among the works it cites.
A unified switching system perspective and ODE analysis of Q-learning algorithms
Donghwan Lee and Niao He · 2019
Later among the works it cites.
Momentum in reinforcement learning
Nino Vieillard, Bruno Scherrer, Olivier Pietquin, and Matthieu Geist · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A finite time analysis of temporal difference learning with linear function approximation
Jalaj Bhandari, Daniel Russo, and Raghav Singal · 2018
Cited alongside, same era.
Q-learning with nearest neighbors
Devavrat Shah and Qiaomin Xie · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Neural temporal-difference learning converges to global optima
Qi Cai, Zhuoran Yang, Jason D Lee, and Zhaoran Wang · 2019
Cited alongside, same era.
Reinforcement learning meets hybrid zero dynamics: A case study for rabbit
Guillermo A Castillo, Bowen Weng, Ayonga Hereid, Zheng Wang, and Wei Zhang · 2019
Cited alongside, same era.
Finite-time analysis of Q-learning with linear function approximation
Zaiwei Chen, Sheng Zhang, Thinh T. Doan, Siva Theja Maguluri, and John-Paul Clarke · 2019
Cited alongside, same era.
Martin J Wainwright · 2019
Later among the works it cites.
A finite-time analysis of Q-learning with neural network function approximation
Pan Xu and Quanquan Gu · 2019
Later among the works it cites.
Sample-optimal parametric Q-learning using linearly additive features
Lin Yang and Mengdi Wang · 2019
Later among the works it cites.
Finite-sample analysis for SARSA with linear function approximation
Shaofeng Zou, Tengyu Xu, and Yingbin Liang · 2019
Later among the works it cites.
Donghwan Lee and Niao He · 2020
Closest in time.
Sample complexity of asynchronous q-learning: Sharper analysis and variance reduction
Gen Li, Yuting Wei, Yuejie Chi, Yuantao Gu, and Yuxin Chen · 2020
Closest in time.
Finite-time analysis of asynchronous stochastic approximation and Q-learning
Guannan Qu and Adam Wierman · 2020
Closest in time.
Analysis of Q-learning with adaptation and momentum restart for gradient descent
Bowen Weng, Huaqing Xiong, Yingbin Liang, and Wei Zhang · 2020
Closest in time.
Non-asymptotic convergence of Adam-type reinforcement learning algorithms under markovian sampling
Huaqing Xiong, Tengyu Xu, Yingbin Liang, and Wei Zhang · 2020
Closest in time.