Fetching the paper…
Reading the bibliography…
We give a simple local Polyak-Lojasiewicz (PL) criterion that guarantees linear (exponential) convergence of gradient flow and gradient descent to a zero-loss solution of a nonnegative objective.
A topological property of real analytic subsets
S. Łojasiewicz · 1963
Earlier work this paper cites.
Gradient methods for minimizing functionals
B. T. Polyak · 1963
Earlier work this paper cites.
Some NP-complete problems in quadratic and nonlinear programming
K. G. Murty and S. N. Kabadi · 1987
Earlier work this paper cites.
On gradients of functions definable in o-minimal structures
K. Kurdyka · 1998
Earlier work this paper cites.
Metric regularity and subdifferential calculus
A. D. Ioffe · 2000
Earlier work this paper cites.
Convex Optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
Y. Nesterov · 2004
Earlier work this paper cites.
Convergence of the iterates of descent methods for analytic cost functions
P.-A. Absil, R. Mahony, and B. Andrews · 2005
Earlier work this paper cites.
Nonlinear error bounds for lower semicontinuous functions on metric spaces
J.-N. Corvellec and V. V. Motreanu · 2008
Earlier work this paper cites.
Proximal alternating minimization and projection methods for nonconvex problems: An approach based on the Kurdyka–Łojasiewicz inequality
H. Attouch, J. Bolte, P. Redont, and A. Soubeyran · 2010
Earlier work this paper cites.
Characterizations of łojasiewicz inequalities: subgradient flows, talweg, convexity
J. Bolte, A. Daniilidis, O. Ley, and L. Mazet · 2010
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
E. Moulines and F. Bach · 2011
Earlier work this paper cites.
Non-strongly-convex smooth stochastic approximation with convergence rate O ( 1 / n ) O(1/n)
F. Bach and E. Moulines · 2013
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
I. J. Goodfellow, O. Vinyals, and A. M. Saxe · 2014
Earlier work this paper cites.
Curves of descent
D. Drusvyatskiy, A. D. Ioffe, and A. S. Lewis · 2015
Earlier work this paper cites.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Earlier work this paper cites.
Identity matters in deep learning
M. Hardt and T. Ma · 2016
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the Polyak–łojasiewicz condition
H. Karimi, J. Nutini, and M. Schmidt · 2016
Cited alongside, same era.
An overview of gradient descent optimization algorithms
S. Ruder · 2016
Cited alongside, same era.
Breaking the curse of dimensionality with convex neural networks
F. Bach · 2017
Cited alongside, same era.
How to escape saddle points efficiently
C. Jin, R. Ge, P. Netrapalli, S. M. Kakade, and M. I. Jordan · 2017
Cited alongside, same era.
A convergence analysis of gradient descent for deep linear neural networks
Lower bounds for non-convex stochastic optimization
Y. Arjevani, Y. Carmon, J. C. Duchi, D. J. Foster, N. Srebro, and B. Woodworth · 2019
Later among the works it cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Y. Cao and Q. Gu · 2019
Later among the works it cites.
On lazy training in differentiable programming
L. Chizat, E. Oyallon, and F. Bach · 2019
Later among the works it cites.
Gradient descent finds global minima of deep neural networks
S. S. Du, J. Lee, H. Li, L. Wang, and X. Zhai · 2019
Later among the works it cites.
Stochastic gradient descent and its variants in machine learning
P. Netrapalli · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Arora, N. Cohen, N. Golowich, and W. Hu · 2018
Cited alongside, same era.
Gradient descent with identity initialization efficiently learns positive definite linear transformations by deep residual networks
P. Bartlett, D. Helmbold, and P. Long · 2018
Cited alongside, same era.
Accelerated methods for nonconvex optimization
Y. Carmon, J. C. Duchi, O. Hinder, and A. Sidford · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
S. S. Du, X. Zhai, B. Poczos, and A. Singh · 2018
Cited alongside, same era.
Stochastic methods for composite and weakly convex optimization problems
J. C. Duchi and F. Ruan · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Cited alongside, same era.
Gradient descent aligns the layers of deep linear networks
Z. Ji and M. Telgarsky · 2018
Cited alongside, same era.
R. Sun · 2019
Later among the works it cites.
Generalization error bounds of gradient descent for learning over-parameterized deep ReLU networks
Y. Cao and Q. Gu · 2020
Later among the works it cites.
Recent theoretical advances in non-convex optimization
M. Danilova, P. Dvurechensky, A. Gasnikov, E. Gorbunov, S. Guminov, D. Kamzolov, and I. Shibaev · 2020
Later among the works it cites.
A comparative analysis of optimization and generalization properties of two-layer neural network and random feature models under gradient descent dynamics
W. E, C. Ma, and L. Wu · 2020
Later among the works it cites.
The impact of neural network overparameterization on gradient confusion and stochastic gradient descent
K. A. Sankararaman, S. De, Z. Xu, W. R. Huang, and T. Goldstein · 2020
Later among the works it cites.
Gradient descent optimizes over-parameterized deep ReLU networks
D. Zou, Y. Cao, D. Zhou, and Q. Gu · 2020
Later among the works it cites.
A. Jentzen and A. Riekert · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2021
Later among the works it cites.
Loss landscapes and optimization in over-parameterized non-linear systems and neural networks
C. Liu, L. Zhu, and M. Belkin · 2022
Closest in time.
Convergence beyond the over-parameterized regime using rayleigh quotients
S. Robin and M. Lelarge · 2022
Closest in time.