Fetching the paper…
Reading the bibliography…
We consider the dynamics of a linear stochastic approximation algorithm driven by Markovian noise, and derive finite-time bounds on the moments of the error, i.e., deviation of the output of the algorithm from the equilibrium point of an associated ordinary differential equation (ODE).
Stochastic approximation methods for decentralized control of multiaccess communications
B. Hajek · 1985
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
R. S. Sutton · 1988
Earlier work this paper cites.
Neuro-dynamic programming
D. P. Bertsekas and J. N. Tsitsiklis · 1996
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
J. N. Tsitsiklis and B. Van Roy · 1997
Earlier work this paper cites.
Linear system theory and design
C. T. Chen · 1998
Earlier work this paper cites.
The ODE method for convergence of stochastic approximation and reinforcement learning
V. S. Borkar and S. P. Meyn · 2000
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications , volume 35
H. Kushner and G. G. Yin · 2003
Earlier work this paper cites.
Stochastic approximation: a dynamical systems viewpoint
V. S. Borkar · 2009
Cited alongside, same era.
Algorithms for reinforcement learning
C. Szepesvári · 2010
Cited alongside, same era.
Dynamic programming and optimal control 3rd edition, volume II
D. P. Bertsekas · 2011
Cited alongside, same era.
Error bounds for constant step-size Q-learning
C. L. Beck and R Srikant · 2012
Cited alongside, same era.
Adaptive algorithms and stochastic approximations , volume 22
A. Benveniste, M. Métivier, and P. Priouret · 2012
Cited alongside, same era.
Stochastic recursive algorithms for optimization: simultaneous perturbation methods , volume 434
S. Bhatnagar, H. L. Prasad, and L. A. Prashanth · 2012
Cited alongside, same era.
Markov chains: Gibbs fields, Monte Carlo Simulation, and Queues , volume 31
P. Brémaud · 2013
Later among the works it cites.
Communication networks: an optimization, control, and stochastic networks perspective
R. Srikant and L. Ying · 2013
Later among the works it cites.
On the approximation error of mean-field models
L. Ying · 2016
Later among the works it cites.
A finite time analysis of temporal difference learning with linear function approximation
J. Bhandari, D. Russo, and R. Singal · 2018
Later among the works it cites.
Finite sample analyses for TD(0) with function approximation
G. Dalal, B. Szörényi, G. Thoppe, and S. Mannor · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Asymptotically tight steady-state queue length bounds implied by drift conditions
A. Eryilmaz and R. Srikant · 2012
Cited alongside, same era.
Simplified description of slow Markov walks. Part I
S. M. Meerkov
Cited in the paper.
Simplified description of slow Markov walks. Part II
S. M. Meerkov
Cited in the paper.
Linear stochastic approximation: How far does constant step-size and iterate averaging go?
C. Lakshminarayanan and C. Szepesvari · 2018
Later among the works it cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.