Fetching the paper…
Reading the bibliography…
In this article, we consider convergence of stochastic gradient descent schemes (SGD), including momentum stochastic gradient descent (MSGD), under weak assumptions on the underlying landscape.
The method of steepest descent for non-linear minimization problems
H. B. Curry · 1944
Earlier work this paper cites.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Une propriété topologique des sous-ensembles analytiques réels
S. Łojasiewicz · 1963
Earlier work this paper cites.
Gradient methods for minimizing functionals
B. T. Poljak · 1963
Earlier work this paper cites.
Ensembles semi-analytiques
S. Łojasiewicz · 1965
Earlier work this paper cites.
Pseudogradient adaptation and training algorithms
B. T. Polyak and Ya. Z. Tsypkin · 1973
Earlier work this paper cites.
Almost sure approximations to the Robbins-Monro and Kiefer-Wolfowitz processes with dependent noise
D. Ruppert · 1982
Earlier work this paper cites.
Adaptive algorithms and stochastic approximations
A. Benveniste, M. Métivier, and P. Priouret · 1990
Earlier work this paper cites.
A new method of stochastic approximation type
B. T. Polyak · 1990
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Foundations of stochastic approximation
H. Walk · 1992
Earlier work this paper cites.
Convergence properties of backpropagation for neural nets via theory of stochastic gradient methods. Part 1
A. A. Gaivoronski · 1994
Earlier work this paper cites.
A class of unconstrained minimization methods for neural network training
L. Grippo · 1994
Earlier work this paper cites.
Analysis of an approximate gradient projection method with applications to the backpropagation algorithm
Z.-Q. Luo and Tseng P · 1994
Earlier work this paper cites.
Serial and parallel backpropagation convergence via nonmonotone perturbed minimization
O. L. Mangasarian and M. V. Solodov · 1994
Earlier work this paper cites.
Algorithmes stochastiques
M. Duflo · 1996
Earlier work this paper cites.
Numerical optimization
J. Nocedal and S. J. Wright · 1999
Earlier work this paper cites.
Gradient convergence in gradient methods with errors
D. P. Bertsekas and J. N. Tsitsiklis · 2000
Earlier work this paper cites.
A primer of real analytic functions
S. G. Krantz and H. R. Parks · 2002
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications
H. J. Kushner and G. G. Yin · 2003
Earlier work this paper cites.
Convergence of the iterates of descent methods for analytic cost functions
P.-A. Absil, R. Mahony, and B. Andrews · 2005
Earlier work this paper cites.
On the convergence of the proximal algorithm for nonsmooth functions involving analytic features
H. Attouch and J. Bolte · 2009
Earlier work this paper cites.
Applications of the Łojasiewicz-Simon gradient inequality to gradient-like evolution equations
R. Chill, A. Haraux, and M. A. Jendoubi · 2009
Earlier work this paper cites.
V. B. Tadic · 2009
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
F. Bach and E. Moulines · 2011
Earlier work this paper cites.
Some applications of the Łojasiewicz gradient inequality
A. Haraux · 2012
Cited alongside, same era.
Martingale in diskreter Zeit: Theorie und Anwendungen
H. Luschgy · 2012
Cited alongside, same era.
Geometric theory of dynamical systems: An introduction
J. Jr Palis and W. De Melo · 2012
Cited alongside, same era.
Non-strongly-convex smooth stochastic approximation with convergence rate O(1/n)
F. Bach and E. Moulines · 2013
Cited alongside, same era.
Convergence and convergence rate of stochastic gradient search in the case of multiple and non-isolated extrema
V. B. Tadic · 2015
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the Polyak-Łojasiewicz condition
H. Karimi, J. Nutini, and M. Schmidt · 2016
Cited alongside, same era.
Reducing parameter space for neural network training
T. Qin, L. Zhou, and D. Xiu · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
B. Woodworth, S. Gunasekar, J. D. Lee, E. Moroshko, P. Savarese, I. Golan, D. Soudry, and N. Srebro · 2020
Later among the works it cites.
Linear convergence of adaptive stochastic gradient descent
Y. Xie, X. Wu, and R. Ward · 2020
Later among the works it cites.
Non-convergence of stochastic gradient descent in the training of deep neural networks
P. Cheridito, A. Jentzen, and F. Rossmannek · 2021
Closest in time.
Global minima of overparameterized neural networks
Y. Cooper · 2021
Closest in time.
SGD for structured nonconvex functions: Learning rates, minibatching and interpolation
R. Gower, O. Sebbouh, and N. Loizou · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Implicit regularization in matrix factorization
S. Gunasekar, B. E. Woodworth, S. Bhojanapalli, B. Neyshabur, and N. Srebro · 2017
Cited alongside, same era.
On exponential convergence of SGD in non-convex over-parametrized learning
R. Bassily, M. Belkin, and S. Ma · 2018
Cited alongside, same era.
Optimal approximation of piecewise smooth functions using deep ReLU neural networks
P. Petersen and F. Voigtlaender · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
D. Soudry, E. Hoffer, M. S. Nacson, S. Gunasekar, and N. Srebro · 2018
Cited alongside, same era.
How SGD selects the global minima in over-parameterized learning: A dynamical stability perspective
L. Wu, C. Ma, and W. E · 2018
Cited alongside, same era.
Implicit regularization in deep matrix factorization
S. Arora, N. Cohen, W. Hu, and Y. Luo · 2019
Cited alongside, same era.
The heavy-tail phenomenon in SGD
M. Gurbuzbalaban, U. Simsekli, and L. Zhu · 2021
Closest in time.
Stochastic gradient descent with noise of machine learning type. Part II: Continuous time analysis
S. Wojtowytsch · 2021
Closest in time.
Convergence rates of the heavy-ball method under the Łojasiewicz property
J.-F. Aujol, C. Dossal, and A. Rondepierre · 2022
Closest in time.
Cooling down stochastic differential equations: Almost sure convergence
S. Dereich and S. Kassing · 2022
Closest in time.
Sharp analysis of stochastic optimization under global Kurdyka-Łojasiewicz inequality
I. Fatkhullin, J. Etesami, N. He, and N. Kiyavash · 2022
Closest in time.
On the existence of global minima and convergence analyses for gradient descent methods in the training of deep neural networks
A. Jentzen and A. Riekert · 2022
Closest in time.
A unified convergence theorem for stochastic optimization methods
X. Li and A. Milzarek · 2022
Closest in time.
A Kurdyka-Łojasiewicz property for stochastic optimization algorithms in a non-convex setting
E. Chouzenoux, J.-B. Fest, and A. Repetti · 2023
Closest in time.
Central limit theorems for stochastic gradient descent with averaging for stable manifolds
S. Dereich and S. Kassing · 2023
Closest in time.
On the existence of optimal shallow feedforward networks with ReLU activation
S. Dereich and S. Kassing · 2023
Closest in time.
Convergence rates for momentum stochastic gradient descent with noise of machine learning type
B. Gess and S. Kassing · 2023
Closest in time.
Overall error analysis for the training of deep neural networks via stochastic gradient descent with random initialisation
A. Jentzen and T. Welti · 2023
Closest in time.
A new inexact gradient descent method with applications to nonsmooth convex optimization
P. D. Khanh, B. S. Mordukhovich, and D. B. Tran · 2023
Closest in time.
Better theory for SGD in the nonconvex world
A. Khaled and P. Richtárik · 2023
Closest in time.
Convergence of random reshuffling under the Kurdyka–Łojasiewicz inequality
X. Li, A. Milzarek, and J. Qiu · 2023
Closest in time.
Convergence of a normal map-based Prox-SGD method under the KL inequality
A. Milzarek and J. Qiu · 2023
Closest in time.
Fast convergence to non-isolated minima: four equivalent conditions for C 2 {C}^{2} functions
Q. Rebjock and N. Boumal · 2023
Closest in time.
Stochastic gradient descent with noise of machine learning type Part I: Discrete time analysis
S. Wojtowytsch · 2023
Closest in time.