Fetching the paper…
Reading the bibliography…
In this article we study fully-connected feedforward deep ReLU ANNs with an arbitrarily large number of hidden layers and we prove convergence of the risk of the GD optimization method with random initializations in the training of such ANNs under the assumption that the unnormalized probability density function of the probability distribution of the input data of the considered supervised learning problem is piecewise polynomial, under the assumption that the target function (describing the relationship between input data and the output data) is piecewise polynomial, and under the assumption that the risk function of the considered supervised learning problem admits at least one regular global minimum.
First-order methods almost always avoid saddle points: the case of vanishing step-sizes, 2019
Ioannis Panageas, Georgios Piliouras, and Xiao Wang · 1906
Earlier work this paper cites.
Deep network approximation characterized by number of neurons, 2020
Zuowei Shen, Haizhao Yang, and Shijun Zhang · 1906
Earlier work this paper cites.
Sur les trajectoires du gradient d’une fonction analytique
S. Łojasiewicz · 1984
Earlier work this paper cites.
Real algebraic geometry
Jacek Bochnak, Michel Coste, and Marie-Francoise Roy · 1987
Earlier work this paper cites.
Semianalytic and subanalytic sets
Edward Bierstone and Pierre D. Milman · 1988
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
G. Cybenko · 1989
Earlier work this paper cites.
Networks and the best approximation property
F. Girosi and T. Poggio · 1990
Earlier work this paper cites.
The analysis of linear partial differential operators. I
Lars Hörmander · 1990
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
Kurt Hornik · 1991
Earlier work this paper cites.
Multilayer feedforward networks with a nonpolynomial activation function can approximate any function
Moshe Leshno, Vladimir Ya. Lin, Allan Pinkus, and Shimon Schocken · 1993
Earlier work this paper cites.
Geometric categories and o-minimal structures
Lou van den Dries and Chris Miller · 1996
Earlier work this paper cites.
Variational analysis
R. Tyrrell Rockafellar and Roger J.-B. Wets · 1998
Earlier work this paper cites.
Gradient convergence in gradient methods with errors
Dimitri P. Bertsekas and John N. Tsitsiklis · 2000
Earlier work this paper cites.
An introduction to semialgebraic geometry
Michel Coste · 2000
Earlier work this paper cites.
Best approximation by heaviside perceptron networks
P. Kainen, V. Kůrková, and A. Vogt · 2000
Earlier work this paper cites.
Proof of the gradient conjecture of R. Thom
Krzysztof Kurdyka, Tadeusz Mostowski, and Adam Parusiński · 2000
Earlier work this paper cites.
Introductory lectures on convex optimization
Yurii Nesterov · 2004
Earlier work this paper cites.
Vivak Patel · 2004
Earlier work this paper cites.
Convergence of the iterates of descent methods for analytic cost functions
P.-A. Absil, R. Mahony, and B. Andrews · 2005
Earlier work this paper cites.
The łojasiewicz inequality for nonsmooth subanalytic functions with applications to subgradient dynamical systems
Jérôme Bolte, Aris Daniilidis, and Adrian Lewis · 2006
Earlier work this paper cites.
Geometry of subanalytic and semialgebraic sets
Masahiro Shiota · 2008
Earlier work this paper cites.
On the convergence of the proximal algorithm for nonsmooth functions involving analytic features
Hedy Attouch and Jérôme Bolte · 2009
Earlier work this paper cites.
Weinan E, Chao Ma, Stephan Wojtowytsch, and Lei Wu · 2009
Earlier work this paper cites.
Partial differential equations
Lawrence C. Evans · 2010
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Eric Moulines and Francis Bach · 2011
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan · 2012
Cited alongside, same era.
Convergence of descent methods for semi-algebraic and tame problems: proximal algorithms, forward-backward splitting, and regularized Gauss-Seidel methods
Hedy Attouch, Jérôme Bolte, and Benar Fux Svaiter · 2013
Cited alongside, same era.
Non-strongly-convex smooth stochastic approximation with convergence rate O ( 1 / n ) O(1/n)
Francis Bach and Eric Moulines · 2013
Cited alongside, same era.
Integration of semialgebraic functions and integrated Nash functions
Tobias Kaiser · 2013
Cited alongside, same era.
A block coordinate descent method for regularized multiconvex optimization with applications to nonnegative tensor factorization and completion
Yangyang Xu and Wotao Yin · 2013
Cited alongside, same era.
Stochastic subgradient method converges on tame functions
Damek Davis, Dmitriy Drusvyatskiy, Sham Kakade, and Jason D. Lee · 2020
Later among the works it cites.
A comparative analysis of optimization and generalization properties of two-layer neural network and random feature models under gradient descent dynamics
Weinan E, Chao Ma, and Lei Wu · 2020
Later among the works it cites.
Convergence rates for the stochastic gradient descent method for non-convex objective functions
Benjamin Fehrman, Benjamin Gess, and Arnulf Jentzen · 2020
Later among the works it cites.
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2020
Later among the works it cites.
Gradient descent optimizes over-parameterized deep ReLU networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Probability theory
Achim Klenke · 2014
Cited alongside, same era.
Escaping from saddle points — online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Cited alongside, same era.
Gradient descent only converges to minimizers
Jason D. Lee, Max Simchowitz, Michael I. Jordan, and Benjamin Recht · 2016
Cited alongside, same era.
Gradient Descent Only Converges to Minimizers: Non-Isolated Critical Points and Invariant Regions
Ioannis Panageas and Georgios Piliouras · 2017
Cited alongside, same era.
An overview of gradient descent optimization algorithms, 2017
Sebastian Ruder · 2017
Cited alongside, same era.
Local minima in training of neural networks, 2017
Grzegorz Swirszcz, Wojciech Marian Czarnecki, and Razvan Pascanu · 2017
Cited alongside, same era.
Optimization methods for large-scale machine learning, 2018
Léon Bottou, Frank E. Curtis, and Jorge Nocedal · 2018
Cited alongside, same era.
Sparse optimization on measures with over-parameterized gradient descent
Lénaïc Chizat · 2021
Closest in time.
Convergence of stochastic gradient descent schemes for Lojasiewicz-landscapes, 2021
Steffen Dereich and Sebastian Kassing · 2021
Closest in time.
Cooling down stochastic differential equations: almost sure convergence, 2021
Steffen Dereich and Sebastian Kassing · 2021
Closest in time.
On minimal representations of shallow ReLU networks, 2021
Steffen Dereich and Sebastian Kassing · 2021
Closest in time.
Simon Eberle, Arnulf Jentzen, Adrian Riekert, and Georg S. Weiss · 2021
Closest in time.
Martin Hutzenthaler, Arnulf Jentzen, Katharina Pohl, Adrian Riekert, and Luca Scarpa · 2021
Closest in time.
Arnulf Jentzen and Timo Kröger · 2021
Closest in time.
Strong error analysis for stochastic gradient descent optimization algorithms
Arnulf Jentzen, Benno Kuckuck, Ariel Neufeld, and Philippe von Wurstemberger · 2021
Closest in time.
Arnulf Jentzen and Adrian Riekert · 2021
Closest in time.
Arnulf Jentzen and Adrian Riekert · 2021
Closest in time.
Arnulf Jentzen and Adrian Riekert · 2021
Closest in time.
Convergence of random reshuffling under the Kurdyka-Łojasiewicz inequality, 2021
Xiao Li, Andre Milzarek, and Junwen Qiu · 2021
Closest in time.
Deep network approximation for smooth functions
Jianfeng Lu, Zuowei Shen, Haizhao Yang, and Shijun Zhang · 2021
Closest in time.
Topological Properties of the Set of Functions Generated by Neural Networks of Fixed Size
Philipp Petersen, Mones Raslan, and Felix Voigtlaender · 2021
Closest in time.
Embedding principle: a hierarchical structure of loss landscape of deep neural networks, 2021
Yaoyu Zhang, Yuqing Li, Zhongwang Zhang, Tao Luo, and Zhi-Qin John Xu · 2021
Closest in time.
Embedding principle of loss landscape of deep neural networks, 2021
Yaoyu Zhang, Zhongwang Zhang, Tao Luo, and Zhi-Qin John Xu · 2021
Closest in time.
Full error analysis for the training of deep neural networks
Christian Beck, Arnulf Jentzen, and Benno Kuckuck · 2022
Closest in time.
A proof of convergence for gradient descent in the training of artificial neural networks for constant target functions
Patrick Cheridito, Arnulf Jentzen, Adrian Riekert, and Florian Rossmannek · 2022
Closest in time.
Landscape Analysis for Shallow Neural Networks: Complete Classification of Critical Points for Affine Target Functions
Patrick Cheridito, Arnulf Jentzen, and Florian Rossmannek · 2022
Closest in time.
Blow up phenomena for gradient descent optimization methods in the training of artificial neural networks, 2022
Davide Gallon, Arnulf Jentzen, and Felix Lindner · 2022
Closest in time.