Fetching the paper…
Reading the bibliography…
The Zap Q-learning algorithm introduced in this paper is an improvement of Watkins' original algorithm and recent competitors in several respects.
A Newton-Raphson version of the multivariate Robbins-Monro procedure
D. Ruppert · 1985
Earlier work this paper cites.
Efficient estimators from a slowly convergent Robbins-Monro processes
D. Ruppert · 1988
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
R. S. Sutton · 1988
Earlier work this paper cites.
Learning from Delayed Rewards
C. J. C. H. Watkins · 1989
Earlier work this paper cites.
Adaptive algorithms and stochastic approximations
A. Benveniste, M. Métivier, and P. Priouret · 1990
Earlier work this paper cites.
Stochastic approximations for finite-state Markov chains
D.-J. Ma, A. M. Makowski, and A. Shwartz · 1990
Earlier work this paper cites.
A new method of stochastic approximation type
B. T. Polyak · 1990
Earlier work this paper cites.
On the Poisson equation for Markov chains: existence of solutions and parameter dependence
A. Shwartz and A. Makowski · 1991
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Q Q -learning
C. J. C. H. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Markov chains and stochastic stability
S. P. Meyn and R. L. Tweedie · 1993
Earlier work this paper cites.
Computable bounds for convergence rates of Markov chains
S. P. Meyn and R. L. Tweedie · 1994
Earlier work this paper cites.
On-line Q-learning using connectionist systems
G. A. Rummery and M. Niranjan · 1994
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
S. J. Bradtke and A. G. Barto · 1996
Cited alongside, same era.
Computable exponential convergence rates for stochastically ordered Markov processes
R. B. Lund, S. P. Meyn, and R. L. Tweedie · 1996
Cited alongside, same era.
Stochastic approximation algorithms and applications
H. J. Kushner and G. G. Yin · 1997
Cited alongside, same era.
The asymptotic convergence-rate of Q-learning
C. Szepesvári · 1997
Cited alongside, same era.
An analysis of temporal-difference learning with function approximation
J. N. Tsitsiklis and B. Van Roy · 1997
Cited alongside, same era.
The ODE method for convergence of stochastic approximation and reinforcement learning
V. S. Borkar and S. P. Meyn · 1998
Cited alongside, same era.
Convergence rate of linear two-time-scale stochastic approximation
V. R. Konda and J. N. Tsitsiklis · 2004
Later among the works it cites.
A generalized Kalman filter for fixed point approximation and efficient temporal-difference learning
D. Choi and B. Van Roy · 2006
Later among the works it cites.
A note on linear function approximation using random projections
K. Barman and V. S. Borkar · 2008
Later among the works it cites.
Stochastic Approximation: A Dynamical Systems Viewpoint
V. S. Borkar · 2008
Later among the works it cites.
Q-learning and Pontryagin’s minimum principle
P. G. Mehta and S. P. Meyn · 2009
Later among the works it cites.
Algorithms for Reinforcement Learning
C. Szepesvári · 2010
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimal stopping of Markov processes: Hilbert space theory, approximation algorithms, and an application to pricing high-dimensional financial derivatives
J. N. Tsitsiklis and B. Van Roy · 1999
Cited alongside, same era.
Technical update: Least-squares temporal difference learning
J. A. Boyan · 2002
Cited alongside, same era.
Hoeffding’s inequality for uniformly ergodic Markov chains
P. W. Glynn and D. Ormoneit · 2002
Cited alongside, same era.
Actor-critic algorithms
V. V. G. Konda · 2002
Cited alongside, same era.
Learning rates for Q-learning
E. Even-Dar and Y. Mansour · 2003
Cited alongside, same era.
Least squares policy evaluation algorithms with linear function approximation
A. Nedic and D. Bertsekas · 2003
Cited alongside, same era.
Speedy Q-learning
M. G. Azar, R. Munos, M. Ghavamzadeh, and H. Kappen · 2011
Later among the works it cites.
TD-learning with exploration
S. P. Meyn and A. Surana · 2011
Later among the works it cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
E. Moulines and F. R. Bach · 2011
Later among the works it cites.
Dynamic Programming and Optimal Control
D. P. Bertsekas · 2012
Later among the works it cites.
Q-learning and policy iteration algorithms for stochastic shortest path problems
H. Yu and D. P. Bertsekas · 2013
Later among the works it cites.
Concentration inequalities for Markov chains by Marton couplings and spectral methods
D. Paulin · 2015
Later among the works it cites.