Fetching the paper…
Reading the bibliography…
Stochastic gradient algorithms are often unstable when applied to functions that do not have Lipschitz-continuous and/or bounded gradients.
On a stochastic approximation method
K. L. Chung · 1954
Earlier work this paper cites.
A convergence theorem for non negative almost supermartingales and some applications
H. Robbins and D. Siegmund · 1971
Earlier work this paper cites.
Stochastic analog of the conjugant-gradient method
A. M. Gupal and L. G. Bazhenov · 1972
Earlier work this paper cites.
The quasigradient method for the solving of the nonlinear programming problems
E. A. Nurminskii · 1973
Earlier work this paper cites.
Stochastic approximation method with gradient averaging for unconstrained problems
A. Ruszczynski and W. Syski · 1983
Earlier work this paper cites.
Minimization methods for non-differentiable functions
N. Z. Shor · 1985
Earlier work this paper cites.
Introduction to optimization
B. T. Polyak · 1987
Earlier work this paper cites.
A linearization method for nonsmooth stochastic programming problems
A. Ruszczyński · 1987
Earlier work this paper cites.
Stochastic quasigradient methods
Y. Ermoliev · 1988
Earlier work this paper cites.
A new algorithm for stochastic optimization
S. Andradóttir · 1990
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Convex analysis and minimization algorithms
J.-B. Hiriart-Urruty and C. Lemaréchal · 1993
Earlier work this paper cites.
A scaled stochastic approximation algorithm
S. Andradöttir · 1996
Earlier work this paper cites.
On the projected subgradient method for nonsmooth convex optimization in a Hilbert space
Y. I. Alber, A. N. Iusem, and M. V. Solodov · 1998
Earlier work this paper cites.
Stochastic generalized gradient method for nonconvex nonsmooth stochastic optimization
Y. M. Ermol’ev and V. Norkin · 1998
Cited alongside, same era.
Subgradient methods
S. Boyd, L. Xiao, and A. Mutapcic · 2003
Cited alongside, same era.
Learning multiple layers of features from tiny images
A. Krizhevsky, G. Hinton, et al · 2009
Cited alongside, same era.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Cited alongside, same era.
Non-Asymptotic Analysis of Stochastic Approximation Algorithms for Machine Learning
F. Bach and E. Moulines · 2011
Cited alongside, same era.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Cited alongside, same era.
Stochastic (approximate) proximal point methods: Convergence, optimality, and adaptivity
H. Asi and J. C. Duchi · 2019
Later among the works it cites.
A stochastic trust region algorithm based on careful step normalization
F. E. Curtis, K. Scheinberg, and R. Shi · 2019
Later among the works it cites.
Stochastic model-based minimization of weakly convex functions
D. Davis and D. Drusvyatskiy · 2019
Later among the works it cites.
Stochastic algorithms with geometric step decay converge linearly on sharp functions
D. Davis, D. Drusvyatskiy, and V. Charisopoulos · 2019
Later among the works it cites.
Proximally guided stochastic subgradient method for nonsmooth, nonconvex problems
D. Davis and B. Grimmer · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the difficulty of training recurrent neural networks
R. Pascanu, T. Mikolov, and Y. Bengio · 2013
Cited alongside, same era.
Stochastic proximal iteration: a non-asymptotic improvement upon stochastic gradient descent
E. K. Ryu and S. Boyd · 2014
Cited alongside, same era.
Beyond convexity: Stochastic quasi-convex optimization
E. Hazan, K. Levy, and S. Shalev-Shwartz · 2015
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
First-order methods in optimization
A. Beck · 2017
Cited alongside, same era.
Introductory lectures on stochastic optimization
J. C. Duchi · 2018
Cited alongside, same era.
J. Zhang, T. He, S. Sra, and A. Jadbabaie · 2019
Later among the works it cites.
Convergence of adaptive algorithms for weakly convex constrained optimization
A. Alacaoglu, Y. Malitsky, and V. Cevher · 2020
Later among the works it cites.
Momentum improves normalized sgd
A. Cutkosky and H. Mehta · 2020
Later among the works it cites.
A single time-scale stochastic approximation method for nested stochastic optimization
S. Ghadimi, A. Ruszczyński, and M. Wang · 2020
Later among the works it cites.
Stochastic optimization with heavy-tailed noise via accelerated gradient clipping
E. Gorbunov, M. Danilova, and A. Gasnikov · 2020
Later among the works it cites.
First-order and Stochastic Optimization Methods for Machine Learning
G. Lan · 2020
Later among the works it cites.
Convergence of a stochastic gradient method with momentum for non-smooth non-convex optimization
V. V. Mai and M. Johansson · 2020
Later among the works it cites.
Improved analysis of clipping algorithms for non-convex optimization
B. Zhang, J. Jin, C. Fang, and L. Wang · 2020
Later among the works it cites.