Fetching the paper…
Reading the bibliography…
In this paper, we introduce a unified framework for analyzing a large family of Q-learning algorithms, based on switching system perspectives and ODE-based stochastic approximation.
Learning to predict by the methods of temporal differences
Richard S. Sutton · 1988
Earlier work this paper cites.
Q-learning
Christopher J. C. H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Convergence of stochastic iterative dynamic programming algorithms
Tommi Jaakkola, Michael I Jordan, and Satinder P Singh · 1994
Earlier work this paper cites.
Asynchronous stochastic approximation and Q-learning
John N Tsitsiklis · 1994
Earlier work this paper cites.
Neuro-dynamic programming
Dimitri P. Bertsekas and John N. Tsitsiklis · 1996
Earlier work this paper cites.
An analog scheme for fixed point computation. i. theory
Vivek S Borkar and K Soumyanatha · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
The asymptotic convergence-rate of Q-learning
Csaba Szepesvári · 1998
Earlier work this paper cites.
Ordinary differential equations (graduate texts in mathematics)
Wolfgang Walter · 1998
Earlier work this paper cites.
The ode method for convergence of stochastic approximation and reinforcement learning
Vivek S Borkar and Sean P Meyn · 2000
Earlier work this paper cites.
Nonlinear systems
Hassan K Khalil · 2002
Cited alongside, same era.
Learning rates for q-learning
Eyal Even-Dar and Yishay Mansour · 2003
Cited alongside, same era.
Stochastic approximation and recursive algorithms and applications , volume 35
Harold Kushner and G. George Yin · 2003
Cited alongside, same era.
Switching in systems and control
Daniel Liberzon · 2003
Cited alongside, same era.
Monotone dynamical systems
Morris W Hirsch and Hal Smith · 2006
Cited alongside, same era.
An analysis of reinforcement learning with function approximation
Francisco S Melo, Sean P Meyn, and M Isabel Ribeiro · 2008
Cited alongside, same era.
Stability and stabilizability of switched linear systems: a survey of recent results
Speedy q-learning
Mohammad Gheshlaghi Azar, Remi Munos, Mohammad Ghavamzadeh, and Hilbert J Kappen · 2011
Later among the works it cites.
Stochastic recursive algorithms for optimization: simultaneous perturbation methods , volume 434
Shalabh Bhatnagar, H. L. Prasad, and L. A. Prashanth · 2012
Later among the works it cites.
Vector barrier certificates and comparison systems
André Platzer · 2018
Later among the works it cites.
Performance of Q-learning with linear function approximation: stability and finite-time analysis, 2019
Zaiwei Chen, Sheng Zhang, Thinh T. Doan, Siva Theja Maguluri, and John-Paul Clarke · 2019
Closest in time.
Adaptive learning rate selection for temporal difference learning
Harsh Gupta, R Srikant, and Lei Ying · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hai Lin and Panos J Antsaklis · 2009
Cited alongside, same era.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Richard S Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvári, and Eric Wiewiora · 2009
Cited alongside, same era.
Double Q-learning
Hado V Hasselt · 2010
Cited alongside, same era.
Toward off-policy learning control with function approximation
Hamid Reza Maei, Csaba Szepesvári, Shalabh Bhatnagar, and Richard S Sutton · 2010
Cited alongside, same era.
Bin Hu and Usman Ahmed Syed · 2019
Closest in time.
Target-based temporal-difference learning
Donghwan Lee and Niao He · 2019
Closest in time.
Finite-time error bounds for linear stochastic approximation and td learning
R Srikant and Lei Ying · 2019
Closest in time.
A theoretical analysis of deep Q-learning
Zhuora Yang, Yuchen Xie, and Zhaoran Wang · 2019
Closest in time.
Finite-sample analysis for sarsa and q-learning with linear function approximation
Shaofeng Zou, Tengyu Xu, and Yingbin Liang · 2019
Closest in time.