Fetching the paper…
Reading the bibliography…
Gradient-based temporal difference (GTD) algorithms are widely used in off-policy learning scenarios.
Learning to predict by the methods of temporal differences
R. S. Sutton · 1988
Earlier work this paper cites.
Q-learning
C. J. Watkins and P. Dayan · 1992
Earlier work this paper cites.
On the generation of Markov decision processes
T. Archibald, K. McKinnon, and L. Thomas · 1995
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
L. Baird · 1995
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
J. N. Tsitsiklis and B. Van Roy · 1997
Earlier work this paper cites.
The ODE method for convergence of stochastic approximation and reinforcement learning
V. S. Borkar and S. P. Meyn · 2000
Earlier work this paper cites.
On the convergence of temporal-difference learning with linear function approximation
V. Tadić · 2001
Earlier work this paper cites.
Convergence rate of linear two-time-scale stochastic approximation
V. R. Konda, J. N. Tsitsiklis, et al · 2004
Earlier work this paper cites.
Almost sure convergence of two time-scale stochastic approximation algorithms
V. B. Tadic · 2004
Earlier work this paper cites.
Convergence rate and averaging of nonlinear two-time-scale stochastic approximation algorithms
A. Mokkadem and M. Pelletier · 2006
Earlier work this paper cites.
A convergent o(n) algorithm for off-policy temporal-difference learning with linear function approximation
R. S. Sutton, C. Szepesvári, and H. R. Maei · 2008
Earlier work this paper cites.
Stochastic approximation: a dynamical systems viewpoint
V. S. Borkar · 2009
Earlier work this paper cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
R. S. Sutton, H. R. Maei, D. Precup, S. Bhatnagar, D. Silver, C. Szepesvári, and E. Wiewiora · 2009
Cited alongside, same era.
On the convergence, lock-in probability and sample complexity of stochastic approximation
S. Kamal · 2010
Cited alongside, same era.
GQ (lambda): A general gradient algorithm for temporal-difference prediction learning with eligibility traces
H. R. Maei and R. S. Sutton · 2010
Cited alongside, same era.
Gradient temporal-difference learning algorithms
H. R. Maei · 2011
Cited alongside, same era.
Policy evaluation with temporal differences: A survey and comparison
C. Dann, G. Neumann, and J. Peters · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
On convergence of some gradient-based temporal-differences algorithms for off-policy learning
H. Yu · 2017
Later among the works it cites.
A finite time analysis of temporal difference learning with linear function approximation
J. Bhandari, D. Russo, and R. Singal · 2018
Later among the works it cites.
Concentration bounds for two time scale stochastic approximation
V. S. Borkar and S. Pattathil · 2018
Later among the works it cites.
Finite sample analyses for TD (0) with function approximation
G. Dalal, B. Szörényi, G. Thoppe, and S. Mannor · 2018
Later among the works it cites.
Finite sample analysis of two-timescale stochastic approximation with applications to reinforcement learning
G. Dalal, B. Szorenyi, G. Thoppe, and S. Mannor · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Finite-sample analysis of proximal gradient td algorithms
B. Liu, J. Liu, M. Ghavamzadeh, S. Mahadevan, and M. Petrik · 2015
Cited alongside, same era.
P. Karmakar and S. Bhatnagar · 2016
Cited alongside, same era.
Stochastic recursive inclusions in two timescales with non-additive iterate dependent Markov noise
V. Yaji and S. Bhatnagar · 2016
Cited alongside, same era.
Two time-scale stochastic approximation with controlled Markov noise and off-policy temporal-difference learning
P. Karmakar and S. Bhatnagar · 2017
Cited alongside, same era.
Finite sample analysis of the GTD policy evaluation algorithms in Markov setting
Y. Wang, W. Chen, Y. Liu, Z.-M. Ma, and T.-Y. Liu · 2017
Cited alongside, same era.
H. R. Maei · 2018
Later among the works it cites.
Stability of stochastic approximations with ’controlled Markov’ noise and temporal difference learning
A. Ramaswamy and S. Bhatnagar · 2018
Later among the works it cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
Why does stagewise training accelerate convergence of testing error over SGD?
T. Yang, Y. Yan, Z. Yuan, and R. Jin · 2018
Later among the works it cites.
Finite-time error bounds for linear stochastic approximation and TD learning
L. Y. R. Srikant · 2019
Closest in time.
A concentration bound for stochastic approximation via Alekseev’s formula
G. Thoppe and V. Borkar · 2019
Closest in time.