Fetching the paper…
Reading the bibliography…
Two-time-scale stochastic approximation, a generalized version of the popular stochastic approximation, has found broad applications in many areas including stochastic control, optimization, and machine learning.
H. Robbins and S. Monro, “A stochastic approximation method,” The Annals of Mathematical Statistics , vol. 22, no. 3, pp. 400–407, 1951
1951
Earlier work this paper cites.
H. Robbins and D. Siegmund, “A convergence theorem for nonnegative almost supermartingales and some applications,” Optimization Methods in Statistics, Academic Press, New York , pp. 233–257, 1971
1971
Earlier work this paper cites.
B. Karimi, B. Miasojedow, E. Moulines, and H.-T. Wai, “Non-asymptotic analysis of biased stochastic approximation scheme,” in Conference on Learning Theory, COLT 2019, 25-28 June 2019, Phoenix, AZ, USA , 2019, pp. 1944–1974
1974
Earlier work this paper cites.
A. Saberi and H. Khalil, “Quadratic-type lyapunov functions for singularly perturbed systems,” IEEE Transactions on Automatic Control , vol. 29, no. 6, pp. 542–550, 1984
1984
Earlier work this paper cites.
J. Chow and P. Kokotovic, “Time scale modeling of sparse dynamic networks,” IEEE Transactions on Automatic Control , vol. 30, no. 8, pp. 714–722, 1985
1985
Earlier work this paper cites.
D. Ruppert, “Efficient estimations from a slowly convergent robbins-monro process,” Technical Report 781, School of Operations Research and Industrial Engineering, Cornell Univ. , 02 1988
1988
Earlier work this paper cites.
B. T. Polyak and A. B. Juditsky, “Acceleration of stochastic approximation by averaging,” SIAM Journal on Control and Optimization , vol. 30, no. 4, pp. 838–855, 1992
1992
Earlier work this paper cites.
D. Bertsekas and J. Tsitsiklis, Neuro-Dynamic Programming , 2nd ed. Athena Scientific, Belmont, MA, 1999
1999
Earlier work this paper cites.
P. Kokotović, H. K. Khalil, and J. O’Reilly, Singular Perturbation Methods in Control: Analysis and Design . Society for Industrial and Applied Mathematics, 1999
1999
Earlier work this paper cites.
H. K. Khalil, Nonlinear System , 3rd ed. Upper Saddle River, NJ: Prentice Hall, 2002
2002
Earlier work this paper cites.
V. R. Konda and J. N. Tsitsiklis, “On actor-critic algorithms,” SIAM J. Control Optim. , vol. 42, no. 4, 2003
2003
Earlier work this paper cites.
V. R. Konda and J. N. Tsitsiklis, “Convergence rate of linear two-time-scale stochastic approximation,” The Annals of Applied Probability , vol. 14, no. 2, pp. 796–819, 2004
2004
Earlier work this paper cites.
A. Mokkadem and M. Pelletier, “Convergence rate and averaging of nonlinear two-time-scale stochastic approximation algorithms,” The Annals of Applied Probability , vol. 16, no. 3, pp. 1671–1702, 2006
2006
Earlier work this paper cites.
V. S. Borkar, Stochastic Approximation: A Dynamical Systems Viewpoint . Cambridge University Press, 2008
2008
Earlier work this paper cites.
E. Biyik and M. Arcak, “Area aggregation and time-scale modeling for sparse nonlinear networks,” Systems and Control Letters , vol. 57, no. 2, pp. 142–149, 2008
2008
Earlier work this paper cites.
T. Hastie, R. Tibshirani, and J. Friedman, The Elements of Statistical Learning : Data Mining, Inference, and Prediction . Springer, 2009
2009
Cited alongside, same era.
R. Sutton, H. R. Maei, and C. Szepesvári, “A convergent o(n) temporal-difference algorithm for off-policy learning with linear function approximation,” in Advances in Neural Information Processing Systems 21 , 2009
2009
Cited alongside, same era.
R. Sutton, H. R. Maei, D. Precup, S. Bhatnagar, D. Silver, C. Szepesvári, and E. Wiewiora, “Fast gradient-descent methods for temporal-difference learning with linear function approximation,” in Proceedings of the 26th International Conference On Machine Learning, ICML , vol. 382, 01 2009
2009
Cited alongside, same era.
H. R. Maei, C. Szepesvári, S. Bhatnagar, D. Precup, D. Silver, and R. S. Sutton, “Convergent temporal-difference learning with arbitrary smooth function approximation,” in Proceedings of the 22nd International Conference on Neural Information Processing Systems , 2009, p. 1204–1212
B. Hu and U. Syed, “Characterizing the exact behaviors of temporal difference learning algorithms using markov jump linear system theory,” in Advances in Neural Information Processing Systems 32 , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
T. T. Doan and J. Romberg, “Linear two-time-scale stochastic approximation a finite-time analysis,” in 2019 57th Annual Allerton Conference on Communication, Control, and Computing (Allerton) , 2019, pp. 399–406
2019
Later among the works it cites.
H. Gupta, R. Srikant, and L. Ying, “Finite-time performance bounds and adaptive learning rate selection for two time-scale reinforcement learning,” in Advances in Neural Information Processing Systems , 2019
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2009
Cited alongside, same era.
A. Benveniste, M. Métivier, and P. Priouret, Adaptive algorithms and stochastic approximations . Springer Science & Business Media, 2012, vol. 22
2012
Cited alongside, same era.
D. Romeres, F. Dörfler, and F. Bullo, “Novel results on slow coherency in consensus and power networks,” in Proc. of 2013 European Control Conference , 2013, pp. 742–747
2013
Cited alongside, same era.
A. M. Boker, C. Yuan, F. Wu, and A. Chakrabortty, “Aggregate control of clustered networks with inter-cluster time delays,” in Proc. of 2016 American Control Conference , 2016, pp. 5340–5345
2016
Cited alongside, same era.
M. Wang, E. X. Fang, and H. Liu, “Stochastic compositional gradient descent: algorithms for minimizing compositions of expected-value functions,” Mathematical Programming , vol. 161, no. 1, Jan 2017
2017
Cited alongside, same era.
T. T. Doan, C. L. Beck, and R. Srikant, “On the convergence rate of distributed gradient methods for finite-sum optimization under communication delays,” Proceedings ACM Meas. Anal. Comput. Syst. , vol. 1, no. 2, pp. 37:1–37:27, 2017
2017
Cited alongside, same era.
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction , 2nd ed. MIT Press, Cambridge, MA, 2018
2018
Cited alongside, same era.
L. Bottou, F. Curtis, and J. Nocedal, “Optimization methods for large-scale machine learning,” SIAM Review , vol. 60, no. 2, pp. 223–311, 2018
2018
Cited alongside, same era.
J. Bhandari, D. Russo, and R. Singal, “A finite time analysis of temporal difference learning with linear function approximation,” in COLT , 2018
2018
Cited alongside, same era.
2019
Later among the works it cites.
G. Lan, Lectures on Optimization Methods for Machine Learning . Springer-Nature, 2020
2020
Closest in time.
T. T. Doan, S. T. Maguluri, and J. Romberg, “Convergence rates of distributed gradient methods under random quantization: A stochastic approximation approach,” IEEE on Transactions on Automatic Control , 2020
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
S. Chen, A. Devraj, A. Busic, and S. Meyn, “Explicit mean-square error bounds for monte-carlo and linear stochastic approximation,” ser. Proceedings of Machine Learning Research, vol. 108, 26–28 Aug 2020, pp. 4173–4183
2020
Closest in time.
M. Kaledin, E. Moulines, A. Naumov, V. Tadic, and H.-T. Wai, “Finite time analysis of linear two-timescale stochastic approximation with Markovian noise,” in Proceedings of Thirty Third Conference on Learning Theory , vol. 125, 2020, pp. 2144–2203
2020
Closest in time.
G. Dalal, B. Szorenyi, and G. Thoppe, “A tale of two-timescale reinforcement learning with the tightest finite-time bound,” Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 04, pp. 3701–3708, Apr. 2020
2020
Closest in time.