Fetching the paper…
Reading the bibliography…
The training of neural networks by gradient descent methods is a cornerstone of the deep learning revolution.
Bounds on rates of variable-basis and neural-network approximation
Vera Kurková and Marcello Sanguineti · 2001
Earlier work this paper cites.
The Łojasiewicz inequality for nonsmooth subanalytic functions with applications to subgradient dynamical systems
Jérôme Bolte, Aris Daniilidis, and Adrian Lewis · 2007
Earlier work this paper cites.
Characterizations of Łojasiewicz inequalities: subgradient flows, talweg, convexity
Jérôme Bolte, Aris Daniilidis, Olivier Ley, and Laurent Mazet · 2010
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lenaic Chizat and Francis Bach · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Gradient descent quantizes ReLu network features
Hartmut Maennel, Olivier Bousquet, and Sylvain Gelly · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Earlier work this paper cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2019
Earlier work this paper cites.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Earlier work this paper cites.
On the inductive bias of neural tangent kernels
Alberto Bietti and Julien Mairal · 2019
Earlier work this paper cites.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Earlier work this paper cites.
Sgd on neural networks learns functions of increasing complexity
Dimitris Kalimeris, Gal Kaplun, Preetum Nakkiran, Benjamin Edelman, Tristan Yang, Boaz Barak, and Haofeng Zhang · 2019
Earlier work this paper cites.
Stochastic modified equations and dynamics of stochastic gradient algorithms i: Mathematical foundations
Qianxiao Li, Cheng Tai, and E Weinan · 2019
Earlier work this paper cites.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2019
Cited alongside, same era.
A function space view of bounded norm infinite width ReLu nets: The multivariate case
Greg Ongie, Rebecca Willett, Daniel Soudry, and Nathan Srebro · 2019
Cited alongside, same era.
How do infinite width bounded norm networks look in function space?
Pedro Savarese, Itay Evron, Daniel Soudry, and Nathan Srebro · 2019
Cited alongside, same era.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lenaic Chizat and Francis Bach · 2020
Cited alongside, same era.
Directional convergence and alignment in deep learning
Ziwei Ji and Matus Telgarsky · 2020
Cited alongside, same era.
The effects of mild over-parameterization on the optimization landscape of shallow ReLU neural networks
Itay M Safran, Gilad Yehudai, and Ohad Shamir · 2021
Later among the works it cites.
Mean-field analysis of piecewise linear solutions for wide ReLu networks
Alexander Shevchenko, Vyacheslav Kungurtsev, and Marco Mondelli · 2021
Later among the works it cites.
Implicit regularization in ReLu networks with the square loss
Gal Vardi and Ohad Shamir · 2021
Later among the works it cites.
Yifei Wang and Mert Pilanci · 2021
Later among the works it cites.
Tensor programs iv: Feature learning in infinite-width neural networks
Greg Yang and Edward J Hu · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhiyuan Li, Yuping Luo, and Kaifeng Lyu · 2020
Cited alongside, same era.
The inductive bias of ReLU networks on orthogonally separable data
Mary Phuong and Christoph H Lampert · 2020
Cited alongside, same era.
Mean field analysis of neural networks: A law of large numbers
Justin Sirignano and Konstantinos Spiliopoulos · 2020
Cited alongside, same era.
On the convergence of gradient descent training for two-layer ReLu-networks in the mean field regime
Stephan Wojtowytsch · 2020
Cited alongside, same era.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Cited alongside, same era.
Numerical influence of ReLu’(0) on backpropagation
David Bertoin, Jérôme Bolte, Sébastien Gerchinovitz, and Edouard Pauwels · 2021
Cited alongside, same era.
Simon Eberle, Arnulf Jentzen, Adrian Riekert, and Georg S Weiss · 2021
Cited alongside, same era.
A unifying view on implicit bias in training linear neural networks
Chulhee Yun, Shankar Krishnan, and Hossein Mobahi · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Later among the works it cites.
A local convergence theory for mildly over-parameterized two-layer neural network
Mo Zhou, Rong Ge, and Chi Jin · 2021
Later among the works it cites.
An initial alignment between neural network and target is needed for gradient descent to learn
Emmanuel Abbe, Elisabetta Cornacchia, Jan Hazla, and Christopher Marquis · 2022
Closest in time.
Convergence of gradient descent for deep neural networks
Sourav Chatterjee · 2022
Closest in time.
On feature learning in neural networks with global convergence guarantees
Zhengdao Chen, Eric Vanden-Eijnden, and Joan Bruna · 2022
Closest in time.
Sparsest piecewise-linear regression of one-dimensional data
Thomas Debarre, Quentin Denoyelle, Michael Unser, and Julien Fageot · 2022
Closest in time.
Loss landscapes and optimization in over-parameterized non-linear systems and neural networks
Chaoyue Liu, Libin Zhu, and Mikhail Belkin · 2022
Closest in time.
What kinds of functions do deep neural networks learn? Insights from variational spline theory
Rahul Parhi and Robert D Nowak · 2022
Closest in time.
Learning sparse features can lead to overfitting in neural networks
Leonardo Petrini, Francesco Cagnetta, Eric Vanden-Eijnden, and Matthieu Wyart · 2022
Closest in time.
Trainability and accuracy of artificial neural networks: An interacting particle system approach
Grant Rotskoff and Eric Vanden-Eijnden · 2022
Closest in time.