Fetching the paper…
Reading the bibliography…
This paper concerns error bounds for recursive equations subject to Markovian disturbances.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Stochastic estimation of the maximum of a regression function
J. Kiefer and J. Wolfowitz · 1952
Earlier work this paper cites.
Multidimensional stochastic approximation methods
J. R. Blum · 1954
Earlier work this paper cites.
On a stochastic approximation method
K. L. Chung et al · 1954
Earlier work this paper cites.
An extension of the robbins-monro procedure
J. Venter et al · 1967
Earlier work this paper cites.
Applications of a Kushner and Clark lemma to general classes of stochastic algorithms
M. Metivier and P. Priouret · 1984
Earlier work this paper cites.
A Newton-Raphson version of the multivariate Robbins-Monro procedure
D. Ruppert · 1985
Earlier work this paper cites.
Efficient estimators from a slowly convergent Robbins-Monro processes
D. Ruppert · 1988
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
R. S. Sutton · 1988
Earlier work this paper cites.
Learning from Delayed Rewards
C. J. C. H. Watkins · 1989
Earlier work this paper cites.
Adaptive algorithms and stochastic approximations
A. Benveniste, M. Métivier, and P. Priouret · 1990
Earlier work this paper cites.
A new method of stochastic approximation type
B. T. Polyak · 1990
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Markov chains and stochastic stability
S. P. Meyn and R. L. Tweedie · 1993
Earlier work this paper cites.
Matrix computations. 3rd. edn. ed, 1996
G. Golub and C. Van Loan · 1996
Cited alongside, same era.
Stochastic approximation algorithms and applications
H. J. Kushner and G. G. Yin · 1997
Cited alongside, same era.
An analysis of temporal-difference learning with function approximation
J. N. Tsitsiklis and B. Van Roy · 1997
Cited alongside, same era.
The ODE method for convergence of stochastic approximation and reinforcement learning
V. S. Borkar and S. P. Meyn · 1998
Cited alongside, same era.
Convergence rate of moments in stochastic approximation with simultaneous perturbation gradient approximation and resetting
L. Gerencser · 1999
Cited alongside, same era.
Hoeffding’s inequality for uniformly ergodic Markov chains
P. W. Glynn and D. Ormoneit · 2002
Cited alongside, same era.
G. Dalal, B. Szorenyi, G. Thoppe, and S. Mannor · 2017
Later among the works it cites.
Fastest convergence for Q-learning
A. M. Devraj and S. P. Meyn · 2017
Later among the works it cites.
Zap Q-learning
A. M. Devraj and S. P. Meyn · 2017
Later among the works it cites.
A finite time analysis of temporal difference learning with linear function approximation
J. Bhandari, D. Russo, and R. Singal · 2018
Later among the works it cites.
Linear stochastic approximation: How far does constant step-size and iterate averaging go?
C. Lakshminarayanan and C. Szepesvari · 2018
Later among the works it cites.
Zap Q Learning with nonlinear function approximation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stochastic approximation and recursive algorithms and applications
H. Kushner and G. G. Yin · 2003
Cited alongside, same era.
Asymptotic analysis of temporal-difference learning algorithms with constant step-sizes
V. B. Tadić · 2006
Cited alongside, same era.
Control Techniques for Complex Networks
S. P. Meyn · 2007
Cited alongside, same era.
Stochastic Approximation: A Dynamical Systems Viewpoint
V. S. Borkar · 2008
Cited alongside, same era.
Most likely paths to error when estimating the mean of a reflected random walk
K. R. Duffy and S. P. Meyn · 2010
Cited alongside, same era.
Adaptive algorithms and stochastic approximations
A. Benveniste, M. Métivier, and P. Priouret · 2012
Cited alongside, same era.
S. Chen, A. M. Devraj, A. Bušić, and S. Meyn · 2019
Later among the works it cites.
Performance of q-learning with linear function approximation: Stability and finite-time analysis
Z. Chen, S. Zhang, T. Doan, S. Maguluri, and J. Clarke · 2019
Later among the works it cites.
Reinforcement Learning Design with Optimal Learning Rate
A. M. Devraj · 2019
Later among the works it cites.
Zap Q-Learning – a user’s guide
A. M. Devraj, A. Bušić, and S. Meyn · 2019
Later among the works it cites.
B. Hu and U. A. Syed · 2019
Later among the works it cites.
Non-asymptotic analysis of biased stochastic approximation scheme
B. Karimi, B. Miasojedow, E. Moulines, and H.-T. Wai · 2019
Later among the works it cites.
Finite-time error bounds for linear stochastic approximation and TD learning
R. Srikant and L. Ying · 2019
Later among the works it cites.
Fundamental design principles for reinforcement learning algorithms
A. M. Devraj, A. Bušić, and S. Meyn · 2020
Closest in time.