Fetching the paper…
Reading the bibliography…
We present for the first time an asymptotic convergence analysis of two time-scale stochastic approximation driven by `controlled' Markov noise.
Principles of Mathematical Analysis
W.Rudin · 1976
Earlier work this paper cites.
Differential Inclusions: Set-Valued Maps and Viability Theory
J.Aubin and A.Cellina · 1984
Earlier work this paper cites.
Applications of a Kushner-Clark lemma to general classes of stochastic algorithms
M.Metivier and P.Priouret · 1984
Earlier work this paper cites.
Adaptive Algorithms and Stochastic Approximation
A.Benveniste, M.Metivier, and P.Priouret · 1990
Earlier work this paper cites.
Stochastic approximations for finite state Markov chains
D.J.Ma, A.M.Makowski, and A.Shwartz · 1990
Earlier work this paper cites.
Probability Theory : An Advanced Course
V.S.Borkar · 1995
Earlier work this paper cites.
Stochastic approximation with two time scales
V.S.Borkar · 1997
Earlier work this paper cites.
Dynamics of stochastic approximation algorithms
M.Benaïm · 1999
Earlier work this paper cites.
Linear stochastic approximation driven by slowly varying Markov chains
V.R.Konda and J.N.Tsitsiklis · 2003
Cited alongside, same era.
On actor-critic algorithms
V.R.Konda and J.N.Tsitsiklis · 2003
Cited alongside, same era.
Almost sure convergence of two time-scale stochastic approximation algorithms
V.B.Tadić · 2004
Cited alongside, same era.
Basis function adaptation in temporal difference reinforcement learning
I.Menache, S.Mannor, and N.Shimkin · 2005
Cited alongside, same era.
Stochastic approximations and differential inclusions
M.Benaïm, J.Hofbauer, and S.Sorin · 2005
Cited alongside, same era.
Stochastic approximation with ‘controlled Markov noise’
V.S.Borkar · 2006
Cited alongside, same era.
Stochastic Approximation : A Dynamic Systems Viewpoint
V.S.Borkar · 2008
Later among the works it cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
R.S.Sutton, H.R.Maei, D.Precup, S.Bhatnagar, D.Silver, and E.Wiewiora · 2009
Later among the works it cites.
Gradient temporal-difference learning algorithms
H.R.Maei · 2011
Later among the works it cites.
Least squares temporal difference methods: an analysis under general conditions
H.Yu · 2012
Later among the works it cites.
Off-policy actor-critic
T.Degris, M.White, and R.S.Sutton · 2012
Later among the works it cites.
Convergence and Convergence Rate of Stochastic Gradient Search in the Case of Multiple and Non-Isolated Extrema
V.B.Tadić · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A convergent O(n) algorithm for off-policy temporal-difference learning with linear function approximation
R.S.Sutton, H.R.Maei, and C.Szepesvári · 2008
Cited alongside, same era.
Weak Convergence Properties of Constrained Emphatic Temporal-difference Learning with Constant and Slowly Diminishing Stepsize
H.Yu · 2016
Closest in time.