Fetching the paper…
Reading the bibliography…
Gradient descent (GD) type optimization schemes are the standard methods to train artificial neural networks (ANNs) with rectified linear unit (ReLU) activation.
Real and complex analysis
Walter Rudin · 1987
Earlier work this paper cites.
Convergence of the iterates of descent methods for analytic cost functions
P.-A. Absil, R. Mahony, and B. Andrews · 2005
Earlier work this paper cites.
The łojasiewicz inequality for nonsmooth subanalytic functions with applications to subgradient dynamical systems
Jérôme Bolte, Aris Daniilidis, and Adrian Lewis · 2006
Earlier work this paper cites.
Weinan E, Chao Ma, Stephan Wojtowytsch, and Lei Wu · 2009
Earlier work this paper cites.
{Euclidean, metric, and Wasserstein} gradient flows: an overview
Filippo Santambrogio · 2017
Earlier work this paper cites.
Solving stochastic differential equations and Kolmogorov equations by means of deep learning, 2018
Christian Beck, Sebastian Becker, Philipp Grohs, Nor Jaafari, and Arnulf Jentzen · 2018
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lénaïc Chizat and Francis Bach · 2018
Earlier work this paper cites.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Simon S Du, Wei Hu, and Jason D Lee · 2018
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks, 2018
Simon S. Du, Xiyu Zhai, Barnabás Poczós, and Aarti Singh · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Cited alongside, same era.
Gradient descent quantizes ReLU network features, 2018
Hartmut Maennel, Olivier Bousquet, and Sylvain Gelly · 2018
Cited alongside, same era.
A geometric approach of gradient descent algorithms in neural networks, 2019
Yacine Chitour, Zhenyu Liao, and Romain Couillet · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Lénaïc Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Convergence rates for the stochastic gradient descent method for non-convex objective functions
Benjamin Fehrman, Benjamin Gess, and Arnulf Jentzen · 2020
Later among the works it cites.
Learning deep linear neural networks: Riemannian gradient flows and convergence to global minimizers
Bubacarr Bah, Holger Rauhut, Ulrich Terstiege, and Michael Westdickenberg · 2021
Closest in time.
Patrick Cheridito, Arnulf Jentzen, Adrian Riekert, and Florian Rossmannek · 2021
Closest in time.
Patrick Cheridito, Arnulf Jentzen, and Florian Rossmannek · 2021
Closest in time.
Sparse optimization on measures with over-parameterized gradient descent
Lénaïc Chizat · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gradient dynamics of shallow univariate ReLU networks
Francis Williams, Matthew Trager, Daniele Panozzo, Claudio Silva, Denis Zorin, and Joan Bruna · 2019
Cited alongside, same era.
A dynamical central limit theorem for shallow neural networks
Zhengdao Chen, Grant Rotskoff, Joan Bruna, and Eric Vanden-Eijnden · 2020
Cited alongside, same era.
A comparative analysis of optimization and generalization properties of two-layer neural network and random feature models under gradient descent dynamics
Weinan E, Chao Ma, and Lei Wu · 2020
Cited alongside, same era.
Closest in time.
Arnulf Jentzen and Adrian Riekert · 2021
Closest in time.
Topological Properties of the Set of Functions Generated by Neural Networks of Fixed Size
Philipp Petersen, Mones Raslan, and Felix Voigtlaender · 2021
Closest in time.