Fetching the paper…
Reading the bibliography…
Motivated by the widespread use of temporal-difference (TD-) and Q-learning algorithms in reinforcement learning, this paper studies a class of biased stochastic approximation (SA) procedures under a mild "ergodic-like" assumption on the underlying stochastic noise sequence.
H. Robbins and S. Monro, “A stochastic approximation method,” Ann. Math. Stat. , vol. 22, no. 3, pp. 400–407, 1951
1951
Earlier work this paper cites.
R. S. Sutton, “Learning to predict by the methods of temporal differences,” Mach. Learn. , vol. 3, no. 1, pp. 9–44, May 1988
1988
Earlier work this paper cites.
C. J. C. H. Watkins, “Learning from delayed rewards,” Ph.D. dissertation, King’s College, Cambridge, 1989
1989
Earlier work this paper cites.
C. J. Watkins and P. Dayan, “Q-learning,” Mach. Learn. , vol. 8, no. 3-4, pp. 279–292, May 1992
1992
Earlier work this paper cites.
T. Jaakkola, M. I. Jordan, and S. P. Singh, “Convergence of stochastic iterative dynamic programming algorithms,” in Adv. in Neural Inf. Process. Syst. , 1994, pp. 703–710
1994
Earlier work this paper cites.
J. N. Tsitsiklis, “Asynchronous stochastic approximation and Q-learning,” Mach. Learn. , vol. 16, no. 3, pp. 185–202, Sept. 1994
1994
Earlier work this paper cites.
L. Baird, “Residual algorithms: Reinforcement learning with function approximation,” in Intl. Conf. on Mach. Learn. , 1995, pp. 30–37
1995
Earlier work this paper cites.
G. J. Gordon, “Stable function approximation in dynamic programming,” in Machine Learning Proceedings . Elsevier, 1995, pp. 261–268
1995
Earlier work this paper cites.
D. P. Bertsekas and J. N. Tsitsiklis, Neuro-Dynamic Programming . Athena Scientific Belmont, MA, 1996, vol. 5
1996
Earlier work this paper cites.
J. N. Tsitsiklis and B. Van Roy, “An analysis of temporal-difference learning with function approximation,” IEEE Trans. Autom. Contr. , vol. 42, no. 5, pp. 674 – 690, May 1997
1997
Earlier work this paper cites.
C. Szepesvári, “The asymptotic convergence-rate of Q-learning,” in Adv. in Neural Inf. Process. Syst. , 1998, pp. 1064–1070
1998
Earlier work this paper cites.
H. Kushner and G. G. Yin, Stochastic Approximation and Recursive Algorithms and Applications . Springer Science & Business Media, 2003, vol. 35
2003
Earlier work this paper cites.
V. S. Borkar, Stochastic Approximation: A Dynamical Systems Viewpoint . Cambridge, New York, NY, 2008, vol. 48
2008
Earlier work this paper cites.
F. S. Melo, S. P. Meyn, and M. I. Ribeiro, “An analysis of reinforcement learning with function approximation,” in Intl. Conf. on Mach. Learn. , 2008, pp. 664–671
2008
Earlier work this paper cites.
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro, “Robust stochastic approximation approach to stochastic programming,” SIAM J. Opt. , vol. 19, no. 4, pp. 1574–1609, Jan. 2009
2009
Cited alongside, same era.
S. Bhatnagar, D. Precup, D. Silver, R. S. Sutton, H. R. Maei, and C. Szepesvári, “Convergent temporal-difference learning with arbitrary smooth function approximation,” in Adv. in Neural Inf. Process. Syst. , 2009, pp. 1204–1212
2009
Cited alongside, same era.
A. Joulin and Y. Ollivier, “Curvature, concentration and error estimates for Markov chain Monte Carlo,” Ann. Prob. , vol. 38, no. 6, pp. 2418–2442, Sep. 2010
2010
Cited alongside, same era.
F. Bach and E. Moulines, “Non-asymptotic analysis of stochastic approximation algorithms for machine learning,” in Adv. in Neural Inf. Process. Syst. , 2011, pp. 451–459
2011
Cited alongside, same era.
2018
Later among the works it cites.
B. Karimi, B. Miasojedow, E. Moulines, and H.-T. Wai, “Non-asymptotic analysis of biased stochastic approximation schemes,” vol. 1, 2019, p. 30
2019
Closest in time.
J. Bhandari, D. Russo, and R. Singal, “A finite time analysis of temporal difference learning with linear function approximation,” in Conf. on Learn. Theory , 2019, pp. 1691–1692
2019
Closest in time.
R. Srikant and L. Ying, “Finite-time error bounds for linear stochastic approximation and TD learning,” 2019
2019
Closest in time.
S. Zou, T. Xu, and Y. Liang, “Finite-sample analysis for SARSA and Q-Learning with linear function approximation,” in Adv. in Neural Inf. Process. Syst. , 2019, pp. 8665–8675
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al. , “Human-level control through deep reinforcement learning,” Nature , vol. 518, no. 7540, p. 529, May 2015
2015
Cited alongside, same era.
N. Korda and P. La, “On TD(0) with function approximation: Concentration bounds and a centered variant with exponential convergence,” in Intl. Conf. on Mach. Learn. , 2015, pp. 626–634
2015
Cited alongside, same era.
B. Liu, J. Liu, M. Ghavamzadeh, S. Mahadevan, and M. Petrik, “Finite-sample analysis of proximal gradient TD algorithms,” in Conf. on Uncertainty in Artif. Intell. , 2015, pp. 504–513
2015
Cited alongside, same era.
N. L. Narayanan and C. Szepesvári, “Finite time bounds for temporal difference learning with function approximation: Problems with some “state-of-the-art” results,” Tech. Rep., 2017
2017
Cited alongside, same era.
D. A. Levin and Y. Peres, Markov Chains and Mixing Times . American Mathematical Society, 2017, vol. 107
2017
Cited alongside, same era.
Y. Wang, W. Chen, Y. Liu, Z.-M. Ma, and T.-Y. Liu, “Finite sample analysis of the GTD policy evaluation algorithms in Markov setting,” in Adv. in Neural Inf. Process. Syst. , 2017, pp. 5504–5513
2017
Cited alongside, same era.
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction . MIT press, 2018
2018
Cited alongside, same era.
G. Dalal, B. Szörényi, G. Thoppe, and S. Mannor, “Finite sample analyses for TD(0) with function approximation,” in AAAI Conf. on Artif. Intell. , 2018, pp. 6144–6152
2018
Cited alongside, same era.
2019
Closest in time.
B. Hu and U. A. Syed, “Characterizing the exact behaviors of temporal difference learning algorithms using Markov jump linear system theory,” in Adv. in Neural Inf. Process. Syst. , 2019, pp. 8477–8488
2019
Closest in time.
Y. Qin, M. Cao, and B. D. O. Anderson, “Lyapunov criterion for stochastic systems and its applications in distributed computation,” IEEE Trans. Autom. Control , pp. 1–15, 2019 (To appear)
2019
Closest in time.
2019
Closest in time.
R. Durrett, Probability: Theory and Examples . Cambridge University Press, 2019, vol. 49
2019
Closest in time.
P. W. Glynn and R. J. Wang, “On the rate of convergence to equilibrium for two-sided reflected Brownian motion and for the Ornstein–Uhlenbeck process,” Queue. Syst. , vol. 91, no. 1-2, pp. 1–14, Feb. 2019
2019
Closest in time.
D. Lee and N. He, “Target-based temporal-difference learning,” in Intl. Conf. on Mach. Learn. , 2019, pp. 3713–3722
2019
Closest in time.
G. Wang, G. B. Giannakis, and J. Chen, “Learning ReLU networks on linearly separable data: Algorithm, optimality, and generalization,” IEEE Trans. Signal Process. , vol. 67, no. 9, pp. 2357–2370, Mar. 2019
2019
Closest in time.