Fetching the paper…
Reading the bibliography…
Stochastic Gradient Descent (SGD) is a widely deployed optimization procedure throughout data-driven and simulation-driven disciplines, which has drawn a substantial interest in understanding its global behavior across a broad class of nonconvex problems and noise models.
A convergence theorem for non negative almost supermartingales and some applications
H. Robbins and D. Siegmund · 1971
Earlier work this paper cites.
Discrete-parameter martingales , volume 10
J. Neveu and T. Speed · 1975
Earlier work this paper cites.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2018
Earlier work this paper cites.
Stochastic (approximate) proximal point methods: Convergence, optimality, and adaptivity
H. Asi and J. C. Duchi · 2019
Cited alongside, same era.
Stochastic gradient descent for nonconvex learning without bounded gradient assumptions
Y. Lei, T. Hu, G. Li, and K. Tang · 2019
Cited alongside, same era.
Sgd for structured nonconvex functions: Learning rates, minibatching and interpolation
R. M. Gower, O. Sebbouh, and N. Loizou · 2020
Cited alongside, same era.
Better theory for sgd in the nonconvex world
A. Khaled and P. Richtárik · 2020
Later among the works it cites.
On the almost sure convergence of stochastic gradient descent in non-convex problems
P. Mertikopoulos, N. Hallak, A. Kavis, and V. Cevher · 2020
Later among the works it cites.
V. Patel · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…