Fetching the paper…
Reading the bibliography…
Stochastic Gradient Descent (SGD) and its variants are the most used algorithms in machine learning applications.
Simple and optimal high-probability bounds for strongly-convex stochastic gradient descent
Harvey, N. J., Liaw, C., and Randhawa, S · 1909
Earlier work this paper cites.
A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Polyak, B. T · 1964
Earlier work this paper cites.
A method for unconstrained convex minimization problem with the rate of convergence O ( 1 / k 2 ) O(1/k^{2})
Nesterov, Y · 1983
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Qian, N · 1999
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87
Nesterov, Y · 2004
Earlier work this paper cites.
On the generalization ability of online strongly convex programming algorithms
Kakade, S. M. and Tewari, A · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Nemirovski, A., Juditsky, A., Lan, G., and Shapiro, A · 2009
Earlier work this paper cites.
Contextual bandit algorithms with supervised learning guarantees
Beygelzimer, A., Langford, J., Li, L., Reyzin, L., and Schapire, R · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J. C., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Validation analysis of mirror descent stochastic approximation method
Lan, G., Nemirovski, A., and Shapiro, A · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
Unified convergence analysis of stochastic momentum methods for convex and non-convex optimization
Yang, T., Lin, Q., and Li, Z · 2016
Cited alongside, same era.
Weighted AdaGrad with unified momentum
Zou, F., Shen, L., Jie, Z., Sun, J., and Liu, W · 2018
Later among the works it cites.
On the convergence of a class of Adam-type algorithms for non-convex optimization
Chen, X., Liu, S., Sun, R., and Hong, M · 2019
Later among the works it cites.
Making the last iterate of SGD information theoretically optimal
Jain, P., Nagaraj, D., and Netrapalli, P · 2019
Later among the works it cites.
A short note on concentration inequalities for random vectors with subgaussian norm
Jin, C., Netrapalli, P., Ge, R., Kakade, S. M., and Jordan, M. I · 2019
Later among the works it cites.
On the convergence of stochastic gradient descent with adaptive stepsizes
Li, X. and Orabona, F · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Loizou, N. and Richtárik, P · 2017
Cited alongside, same era.
Stochastic heavy ball
Gadat, S., Panloup, F., and Saadane, S · 2018
Cited alongside, same era.
On the convergence of Adam and beyond
Reddi, S. J., Kale, S., and Kumar, S · 2018
Cited alongside, same era.
On the convergence of adaptive gradient methods for nonconvex optimization, 2018
Zhou, D., Tang, Y., Yang, Z., Cao, Y., and Gu, Q · 2018
Cited alongside, same era.
Tight analyses for non-smooth stochastic gradient descent
Harvey, N. J., Liaw, C., Plan, Y., and Randhawa, S
Cited in the paper.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Later among the works it cites.
Domain-independent dominance of adaptive methods
Savarese, P., McAllester, D., Babu, S., and Maire, M · 2019
Later among the works it cites.
AdaGrad stepsizes: Sharp convergence over nonconvex landscapes, from any initialization
Ward, R., Wu, X., and Bottou, L · 2019
Later among the works it cites.
A sufficient condition for convergences of Adam and RMSProp
Zou, F., Shen, L., Jie, Z., Zhang, W., and Liu, W · 2019
Later among the works it cites.