Fetching the paper…
Reading the bibliography…
We study the so-called two-time-scale stochastic approximation, a simulation-based approach for finding the roots of two coupled nonlinear operators.
H. Robbins and S. Monro, “A stochastic approximation method,” The Annals of Mathematical Statistics
1951
Earlier work this paper cites.
B. Poljak and J. Tsypkin, “Robust identification,” Automatica
1980
Earlier work this paper cites.
D. Ruppert, “Efficient estimations from a slowly convergent robbins-monro process,” Technical Report 781, School of Operations Research and Industrial Engineering, Cornell Univ
1988
Earlier work this paper cites.
B. T. Polyak and A. B. Juditsky, “Acceleration of stochastic approximation by averaging,” SIAM Journal on Control and Optimization
1992
Earlier work this paper cites.
Society for Industrial and Applied Mathematics, 1999
P. Kokotović, H. K. Khalil, and J. O’Reilly, Singular Perturbation Methods in Control: Analysis and Design · 1999
Earlier work this paper cites.
V. R. Konda and J. N. Tsitsiklis, “On actor-critic algorithms,” SIAM J. Control Optim
2003
Earlier work this paper cites.
C. Andrieu, N. de Freitas, A. Doucet, and M. I. Jordan, “An introduction to mcmc for machine learning,” Machine Learning
2003
Earlier work this paper cites.
Springer, NY, 2nd ed., 2003
H. Kushner and G. Yin, Stochastic Approximation and Recursive Algorithms and Applications · 2003
Earlier work this paper cites.
V. R. Konda and J. N. Tsitsiklis, “Convergence rate of linear two-time-scale stochastic approximation,” The Annals of Applied Probability
2004
Earlier work this paper cites.
V. S. Borkar, “An actor-critic algorithm for constrained markov decision processes,” Systems & Control Letters
2005
Earlier work this paper cites.
A. Mokkadem and M. Pelletier, “Convergence rate and averaging of nonlinear two-time-scale stochastic approximation algorithms,” The Annals of Applied Probability
2006
Earlier work this paper cites.
American Mathematical Society, 2006
D. A. Levin, Y. Peres, and E. L. Wilmer, Markov chains and mixing times · 2006
Earlier work this paper cites.
Cambridge University Press, 2008
V. S. Borkar, Stochastic Approximation: A Dynamical Systems Viewpoint · 2008
Earlier work this paper cites.
S. Ram, A. Nedić, and V. V. Veeravalli, “Incremental stochastic subgradient algorithms for convex optimization,” SIAM Journal on Optimization
2009
Earlier work this paper cites.
H. R. Maei, C. Szepesvári, S. Bhatnagar, D. Precup, D. Silver, and R. S. Sutton, “Convergent temporal-difference learning with arbitrary smooth function approximation,” in Proceedings of the 22nd International Conference on Neural Information Processing Systems
2009
Earlier work this paper cites.
B. Johansson, M. Rabi, and M. Johansson, “A randomized incremental subgradient method for distributed optimization in networked systems,” SIAM Journal on Optimization
2010
Earlier work this paper cites.
Springer Science & Business Media, 2012
A. Benveniste, M. Métivier, and P. Priouret, Adaptive algorithms and stochastic approximations · 2012
Cited alongside, same era.
M. Wang, E. X. Fang, and H. Liu, “Stochastic compositional gradient descent: algorithms for minimizing compositions of expected-value functions,” Mathematical Programming
2017
Cited alongside, same era.
T. T. Doan, C. L. Beck, and R. Srikant, “On the convergence rate of distributed gradient methods for finite-sum optimization under communication delays,” Proceedings ACM Meas. Anal. Comput. Syst
2017
Cited alongside, same era.
2017
Cited alongside, same era.
MIT Press, Cambridge, MA, 2nd ed., 2018
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction · 2018
Cited alongside, same era.
Springer-Nature, 2020
G. Lan, Lectures on Optimization Methods for Machine Learning · 2020
Later among the works it cites.
2020
Later among the works it cites.
T. T. Doan, S. T. Maguluri, and J. Romberg, “Convergence rates of distributed gradient methods under random quantization: A stochastic approximation approach,” IEEE on Transactions on Automatic Control
2020
Later among the works it cites.
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Sun, Y. Sun, and W. Yin, “On markov chain gradient descent,” in Proceedings of the 32nd International Conference on Neural Information Processing Systems
2018
Cited alongside, same era.
L. Bottou, F. Curtis, and J. Nocedal, “Optimization methods for large-scale machine learning,” SIAM Review
2018
Cited alongside, same era.
G. Dalal, G. Thoppe, B. Szörényi, and S. Mannor, “Finite sample analysis of two-timescale stochastic approximation with applications to reinforcement learning,” in COLT
2018
Cited alongside, same era.
J. Zhang and L. Xiao, “A stochastic composite gradient method with incremental variance reduction,” in Advances in Neural Information Processing Systems 32
2019
Cited alongside, same era.
B. Karimi, B. Miasojedow, E. Moulines, and H.-T. Wai, “Non-asymptotic analysis of biased stochastic approximation scheme,” in Conference on Learning Theory, COLT 2019, 25-28 June 2019, Phoenix, AZ, USA
2019
Cited alongside, same era.
R. Srikant and L. Ying, “Finite-time error bounds for linear stochastic approximation and TD learning,” in COLT
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2020
Later among the works it cites.
Z. Chen, S. T. Maguluri, S. Shakkottai, and K. Shanmugam, “Finite-sample analysis of contractive stochastic approximation using smooth convex envelopes,” in Proceedings of the International Conference on Neural Information Processing Systems
2020
Later among the works it cites.
G. Qu and A. Wierman, “Finite-time analysis of asynchronous stochastic approximation and q q -learning,” in Proceedings of Thirty Third Conference on Learning Theory
2020
Later among the works it cites.
W. Mou, C. J. Li, M. J. Wainwright, P. L. Bartlett, and M. I. Jordan, “On linear stochastic approximation: Fine-grained Polyak-Ruppert and non-asymptotic concentration,” in Proceedings of Thirty Third Conference on Learning Theory
2020
Later among the works it cites.
S. Chen, A. Devraj, A. Busic, and S. Meyn, “Explicit mean-square error bounds for monte-carlo and linear stochastic approximation,” vol. 108 of Proceedings of Machine Learning Research
2020
Later among the works it cites.
2020
Later among the works it cites.
D. Nagaraj, X. Wu, G. Bresler, T. Jain, and P. Netrapalli, “Least squares regression with markovian data: Fundamental limits and algorithms,” in Advances in Neural Information Processing Systems 33
2020
Later among the works it cites.
G. Dalal, B. Szorenyi, and G. Thoppe, “A tale of two-timescale reinforcement learning with the tightest finite-time bound,” Proceedings of the AAAI Conference on Artificial Intelligence
2020
Later among the works it cites.
M. Kaledin, E. Moulines, A. Naumov, V. Tadic, and H.-T. Wai, “Finite time analysis of linear two-timescale stochastic approximation with Markovian noise,” in Proceedings of Thirty Third Conference on Learning Theory
2020
Later among the works it cites.
2020
Later among the works it cites.