Fetching the paper…
Reading the bibliography…
Motivated by their broad applications in reinforcement learning, we study the linear two-time-scale stochastic approximation, an iterative method using two different step sizes for finding the solutions of a system of two equations.
H. Robbins and S. Monro, “A stochastic approximation method,” The Annals of Mathematical Statistics , vol. 22, no. 3, pp. 400–407, 1951
1951
Earlier work this paper cites.
B. Karimi, B. Miasojedow, E. Moulines, and H. Wai, “Non-asymptotic analysis of biased stochastic approximation scheme,” in Conference on Learning Theory, COLT 2019, 25-28 June 2019, Phoenix, AZ, USA , 2019, pp. 1944–1974
1974
Earlier work this paper cites.
R. S. Sutton, “Learning to predict by the methods of temporal differences,” Machine Learning , vol. 3, no. 1, pp. 9–44, Aug 1988
1988
Earlier work this paper cites.
B. T. Polyak and A. B. Juditsky, “Acceleration of stochastic approximation by averaging,” SIAM Journal on Control and Optimization , vol. 30, no. 4, pp. 838–855, 1992
1992
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction , 1st ed. MIT Press, 1998
1998
Earlier work this paper cites.
V. Borkar and S. Meyn, “The o.d.e. method for convergence of stochastic approximation and reinforcement learning,” SIAM Journal on Control and Optimization , vol. 38, no. 2, pp. 447–469, 2000
2000
Earlier work this paper cites.
V. R. Konda and J. N. Tsitsiklis, “On actor-critic algorithms,” SIAM J. Control Optim. , vol. 42, no. 4, 2003
2003
Earlier work this paper cites.
V. R. Konda and J. N. Tsitsiklis, “Convergence rate of linear two-time-scale stochastic approximation,” The Annals of Applied Probability , vol. 14, no. 2, pp. 796–819, 2004
2004
Earlier work this paper cites.
V. S. Borkar, “An actor-critic algorithm for constrained markov decision processes,” Systems & Control Letters , vol. 54, no. 3, pp. 207 – 213, 2005
2005
Earlier work this paper cites.
A. Mokkadem and M. Pelletier, “Convergence rate and averaging of nonlinear two-time-scale stochastic approximation algorithms,” The Annals of Applied Probability , vol. 16, no. 3, pp. 1671–1702, 2006
2006
Earlier work this paper cites.
V. Borkar, Stochastic Approximation: A Dynamical Systems Viewpoint . Cambridge University Press, 2008
2008
Earlier work this paper cites.
R. Sutton, H. R. Maei, and C. Szepesvári, “A convergent o(n) temporal-difference algorithm for off-policy learning with linear function approximation,” in Advances in Neural Information Processing Systems 21 , 2009
2009
Cited alongside, same era.
R. Sutton, H. R. Maei, D. Precup, S. Bhatnagar, D. Silver, C. Szepesvári, and E. Wiewiora, “Fast gradient-descent methods for temporal-difference learning with linear function approximation,” in Proceedings of the 26th International Conference On Machine Learning, ICML , vol. 382, 01 2009
2009
Cited alongside, same era.
B. Pierre, Markov Chains: Gibbs Fields, Monte Carlo Simulation, and Queues . Springer Science & Business Media, 01 2013, vol. 31
2013
Cited alongside, same era.
G. Lan, “Gradient sliding for composite optimization,” Math. Program. , vol. 159, no. 1-2, pp. 201–235, Sep. 2016
2016
Cited alongside, same era.
D. Lee and N. He, “Target-based temporal-difference learning,” in Proceedings of the 36th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, vol. 97. Long Beach, California, USA: PMLR, 09–15 Jun 2019, pp. 3713–3722
2019
Closest in time.
2019
Closest in time.
T. Xu, S. Zou, and Y. Liang, “Two time-scale off-policy td learning: Non-asymptotic analysis over markovian samples,” in Advances in Neural Information Processing Systems 32 , 2019
2019
Closest in time.
2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Wang, E. X. Fang, and H. Liu, “Stochastic compositional gradient descent: algorithms for minimizing compositions of expected-value functions,” Mathematical Programming , vol. 161, no. 1, Jan 2017
2017
Cited alongside, same era.
T. T. Doan, C. L. Beck, and R. Srikant, “On the convergence rate of distributed gradient methods for finite-sum optimization under communication delays,” Proceedings ACM Meas. Anal. Comput. Syst. , vol. 1, no. 2, pp. 37:1–37:27, 2017
2017
Cited alongside, same era.
2018
Cited alongside, same era.
L. Bottou, F. Curtis, and J. Nocedal, “Optimization methods for large-scale machine learning,” SIAM Review , vol. 60, no. 2, pp. 223–311, 2018
2018
Cited alongside, same era.
J. Bhandari, D. Russo, and R. Singal, “A finite time analysis of temporal difference learning with linear function approximation,” in COLT , 2018
2018
Cited alongside, same era.
G. Dalal, G. Thoppe, B. Szörényi, and S. Mannor, “Finite sample analysis of two-timescale stochastic approximation with applications to reinforcement learning,” in COLT , 2018
2018
Cited alongside, same era.
R. Srikant and L. Ying, “Finite-time error bounds for linear stochastic approximation and TD learning,” in COLT , 2019
2019
Closest in time.
2019
Closest in time.
B. Hu and U. Syed, “Characterizing the exact behaviors of temporal difference learning algorithms using markov jump linear system theory,” in Advances in Neural Information Processing Systems 32 , 2019
2019
Closest in time.
T. T. Doan and J. Romberg, “Linear two-time-scale stochastic approximation a finite-time analysis,” in 2019 57th Annual Allerton Conference on Communication, Control, and Computing (Allerton) , 2019, pp. 399–406
2019
Closest in time.
H. Gupta, R. Srikant, and L. Ying, “Finite-time performance bounds and adaptive learning rate selection for two time-scale reinforcement learning,” in Advances in Neural Information Processing Systems , 2019
2019
Closest in time.