Fetching the paper…
Reading the bibliography…
Motivated by broad applications in reinforcement learning and machine learning, this paper considers the popular stochastic gradient descent (SGD) when the gradients of the underlying objective function are sampled from Markov processes.
H. Robbins and S. Monro, “A stochastic approximation method,” The Annals of Mathematical Statistics , vol. 22, no. 3, pp. 400–407, 1951
1951
Earlier work this paper cites.
B. Karimi, B. Miasojedow, E. Moulines, and H. Wai, “Non-asymptotic analysis of biased stochastic approximation scheme,” in Conference on Learning Theory, COLT 2019, 25-28 June 2019, Phoenix, AZ, USA , 2019, pp. 1944–1974
1974
Earlier work this paper cites.
Z.-Q. Luo and P. Tseng, “Error bounds and convergence analysis of feasible descent methods: a general approach,” p. 157–178, 1993
1993
Earlier work this paper cites.
S. Ram, A. Nedić, and V. V. Veeravalli, “Incremental stochastic subgradient algorithms for convex optimization,” SIAM Journal on Optimization , vol. 20, no. 2, pp. 691–717, 2009
2009
Earlier work this paper cites.
B. Johansson, M. Rabi, and M. Johansson, “A randomized incremental subgradient method for distributed optimization in networked systems,” SIAM Journal on Optimization , vol. 20, no. 3, pp. 1157–1170, 2010
2010
Earlier work this paper cites.
J. C. Duchi, A. Agarwal, M. Johansson, and M. Jordan, “Ergodic mirror descent,” SIAM Journal on Optimization , vol. 22, no. 4, pp. 1549–1578, 2012
2012
Earlier work this paper cites.
S. Ghadimi and G. Lan, “Stochastic first- and zeroth-order methods for nonconvex stochastic programming,” SIAM Journal on Optimization , vol. 23, no. 4, pp. 2341–2368, 2013
2013
Earlier work this paper cites.
H. Karimi, J. Nutini, and M. Schmidt, “Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition,” in Machine Learning and Knowledge Discovery in Databases . Springer International Publishing, 2016, pp. 795–811
2016
Cited alongside, same era.
F. Schopfer, “Linear convergence of descent methods for the unconstrained minimization of restricted strongly convex functions,” SIAM Journal on Optimization , vol. 26, no. 3, pp. 1883–1911, 2016
2016
Cited alongside, same era.
J. Bolte, T. P. Nguyen, J. Peypouquet, and B. W. Suter, “From error bounds to the complexity of first-order descent methods for convex functions,” pp. 471–507, 2017
2017
Cited alongside, same era.
H. Zhang, “The restricted strong convexity revisited: analysis of equivalence to error bound and quadratic growth,” pp. 817–833, 2017
2017
Cited alongside, same era.
L. M. Nguyen, P. H. Nguyen, P. Richtarik, K. Scheinberg, M. Takac, and M. van Dijk, “New convergence aspects of stochastic gradient algorithms,” Journal of Machine Learning Research , vol. 20, no. 176, pp. 1–49, 2019
2019
Later among the works it cites.
R. Srikant and L. Ying, “Finite-time error bounds for linear stochastic approximation and TD learning,” in COLT , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Bottou, F. E. Curtis, and J. Nocedal, “Optimization methods for large-scale machine learning,” Siam Review , vol. 60, no. 2, pp. 223–311, 2018
2018
Cited alongside, same era.
L. Nguyen, P. H. Nguyen, M. van Dijk, P. Richtarik, K. Scheinberg, and M. Takac, “SGD and Hogwild! convergence without the bounded gradients assumption,” in International Conference on Machine Learning , 2018, pp. 3747–3755
2018
Cited alongside, same era.
T. Sun, Y. Sun, and W. Yin, “On markov chain gradient descent,” in Proceedings of the 32nd International Conference on Neural Information Processing Systems , ser. NIPS’18. Red Hook, NY, USA: Curran Associates Inc., 2018, p. 9918–9927
2018
Cited alongside, same era.
2019
Later among the works it cites.