Fetching the paper…
Reading the bibliography…
The training of artificial neural networks (ANNs) with rectified linear unit (ReLU) activation via gradient descent (GD) type optimization schemes is nowadays a common industrially relevant procedure.
Sur les trajectoires du gradient d’une fonction analytique
S. Łojasiewicz · 1984
Earlier work this paper cites.
Semianalytic and subanalytic sets
Edward Bierstone and Pierre D. Milman · 1988
Earlier work this paper cites.
Variational analysis
R. Tyrrell Rockafellar and Roger J.-B. Wets · 1998
Earlier work this paper cites.
Gradient convergence in gradient methods with errors
Dimitri P. Bertsekas and John N. Tsitsiklis · 2000
Earlier work this paper cites.
An introduction to semialgebraic geometry
Michel Coste · 2000
Earlier work this paper cites.
Proof of the gradient conjecture of R. Thom
Krzysztof Kurdyka, Tadeusz Mostowski, and Adam Parusiński · 2000
Earlier work this paper cites.
Introductory lectures on convex optimization
Yurii Nesterov · 2004
Earlier work this paper cites.
Vivak Patel · 2004
Earlier work this paper cites.
Convergence of the iterates of descent methods for analytic cost functions
P.-A. Absil, R. Mahony, and B. Andrews · 2005
Earlier work this paper cites.
The łojasiewicz inequality for nonsmooth subanalytic functions with applications to subgradient dynamical systems
Jérôme Bolte, Aris Daniilidis, and Adrian Lewis · 2006
Earlier work this paper cites.
On the convergence of the proximal algorithm for nonsmooth functions involving analytic features
Hedy Attouch and Jérôme Bolte · 2009
Earlier work this paper cites.
Weinan E, Chao Ma, Stephan Wojtowytsch, and Lei Wu · 2009
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Eric Moulines and Francis Bach · 2011
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan · 2012
Earlier work this paper cites.
Convergence of descent methods for semi-algebraic and tame problems: proximal algorithms, forward-backward splitting, and regularized Gauss-Seidel methods
Hedy Attouch, Jérôme Bolte, and Benar Fux Svaiter · 2013
Cited alongside, same era.
Non-strongly-convex smooth stochastic approximation with convergence rate O ( 1 / n ) O(1/n)
Francis Bach and Eric Moulines · 2013
Cited alongside, same era.
Integration of semialgebraic functions and integrated Nash functions
Tobias Kaiser · 2013
Cited alongside, same era.
Gradient descent only converges to minimizers
Jason D. Lee, Max Simchowitz, Michael I. Jordan, and Benjamin Recht · 2016
Cited alongside, same era.
{Euclidean, metric, and Wasserstein} gradient flows: an overview
Filippo Santambrogio · 2017
Cited alongside, same era.
On the global convergence of gradient descent for over-parameterized models using optimal transport
A dynamical central limit theorem for shallow neural networks
Zhengdao Chen, Grant Rotskoff, Joan Bruna, and Eric Vanden-Eijnden · 2020
Later among the works it cites.
A comparative analysis of optimization and generalization properties of two-layer neural network and random feature models under gradient descent dynamics
Weinan E, Chao Ma, and Lei Wu · 2020
Later among the works it cites.
Convergence rates for the stochastic gradient descent method for non-convex objective functions
Benjamin Fehrman, Benjamin Gess, and Arnulf Jentzen · 2020
Later among the works it cites.
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lénaïc Chizat and Francis Bach · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Cited alongside, same era.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Lénaïc Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S. Du, Xiyu Zhai, Barnabás Póczos, and Aarti Singh · 2019
Cited alongside, same era.
First-order methods almost always avoid strict saddle points
Jason D. Lee, Ioannis Panageas, Georgios Piliouras, Max Simchowitz, Michael I. Jordan, and Benjamin Recht · 2019
Cited alongside, same era.
Stochastic gradient descent for nonconvex learning without bounded gradient assumptions
Y. Lei, T. Hu, G. Li, and K. Tang · 2019
Cited alongside, same era.
Patrick Cheridito, Arnulf Jentzen, Adrian Riekert, and Florian Rossmannek · 2021
Closest in time.
Patrick Cheridito, Arnulf Jentzen, and Florian Rossmannek · 2021
Closest in time.
Sparse optimization on measures with over-parameterized gradient descent
Lénaïc Chizat · 2021
Closest in time.
Convergence of stochastic gradient descent schemes for Lojasiewicz-landscapes, 2021
Steffen Dereich and Sebastian Kassing · 2021
Closest in time.
Arnulf Jentzen and Timo Kröger · 2021
Closest in time.
Strong error analysis for stochastic gradient descent optimization algorithms
Arnulf Jentzen, Benno Kuckuck, Ariel Neufeld, and Philippe von Wurstemberger · 2021
Closest in time.
Arnulf Jentzen and Adrian Riekert · 2021
Closest in time.
Arnulf Jentzen and Adrian Riekert · 2021
Closest in time.
Arnulf Jentzen and Adrian Riekert · 2021
Closest in time.