Fetching the paper…
Reading the bibliography…
Training neural networks with first order optimisation methods is at the core of the empirical success of deep learning.
Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition
Thomas M Cover · 1965
Earlier work this paper cites.
Optimization and nonsmooth analysis
Frank H Clarke · 1990
Earlier work this paper cites.
Manifolds
Loring W Tu · 2011
Earlier work this paper cites.
Differential inclusions: set-valued maps and viability theory , volume 264
J-P Aubin and Arrigo Cellina · 2012
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Earlier work this paper cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Earlier work this paper cites.
Overfitting or perfect fitting? risk bounds for classification and regression rules that interpolate
Mikhail Belkin, Daniel J Hsu, and Partha Mitra · 2018
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lenaic Chizat and Francis Bach · 2018
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Earlier work this paper cites.
Gradient descent happens in a tiny subspace
Guy Gur-Ari, Daniel A Roberts, and Ethan Dyer · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Gradient descent quantizes relu network features
Hartmut Maennel, Olivier Bousquet, and Sylvain Gelly · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Earlier work this paper cites.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Earlier work this paper cites.
Gradient descent aligns the layers of deep linear networks
Ziwei Ji and Matus Telgarsky · 2019
Earlier work this paper cites.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2019
Earlier work this paper cites.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2020
Earlier work this paper cites.
A mathematical model for automatic differentiation in machine learning
Jérôme Bolte and Edouard Pauwels · 2020
Cited alongside, same era.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lenaic Chizat and Francis Bach · 2020
Cited alongside, same era.
Bad global minima exist and sgd can reach them
Shengchao Liu, Dimitris Papailiopoulos, and Dimitris Achlioptas · 2020
Cited alongside, same era.
The inductive bias of relu networks on orthogonally separable data
Mary Phuong and Christoph H Lampert · 2020
Cited alongside, same era.
Implicit regularization in deep learning may not be explainable by norms
Noam Razin and Nadav Cohen · 2020
Cited alongside, same era.
On the convergence of gradient descent training for two-layer relu-networks in the mean field regime
Gradient flow dynamics of shallow relu networks for square loss and orthogonal inputs
Etienne Boursier, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2022
Later among the works it cites.
Convergence of gradient descent for deep neural networks
Sourav Chatterjee · 2022
Later among the works it cites.
Benign overfitting without linearity: Neural network classifiers trained by gradient descent for noisy linear data
Spencer Frei, Niladri S Chatterji, and Peter Bartlett · 2022
Later among the works it cites.
Trainability and accuracy of artificial neural networks: An interacting particle system approach
Grant Rotskoff and Eric Vanden-Eijnden · 2022
Later among the works it cites.
On the effective number of linear regions in shallow univariate relu networks: Convergence guarantees and implicit bias
Itay Safran, Gal Vardi, and Jason D Lee · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stephan Wojtowytsch · 2020
Cited alongside, same era.
Neural networks as kernel learners: The silent alignment effect
Alexander Atanasov, Blake Bordelon, and Cengiz Pehlevan · 2021
Cited alongside, same era.
Conservative set valued fields, automatic differentiation, stochastic gradient methods and deep learning
Jérôme Bolte and Edouard Pauwels · 2021
Cited alongside, same era.
Stochastic training is not necessary for generalization
Jonas Geiping, Micah Goldblum, Phil Pope, Michael Moeller, and Tom Goldstein · 2021
Cited alongside, same era.
Arthur Jacot, François Ged, Berfin Şimşek, Clément Hongler, and Franck Gabriel · 2021
Cited alongside, same era.
Highly accurate protein structure prediction with alphafold
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al · 2021
Cited alongside, same era.
Gradient descent on two-layer nets: Margin maximization and simplicity bias
Kaifeng Lyu, Zhiyuan Li, Runzhe Wang, and Sanjeev Arora · 2021
Cited alongside, same era.
Mean-field analysis of piecewise linear solutions for wide relu networks
Alexander Shevchenko, Vyacheslav Kungurtsev, and Marco Mondelli · 2022
Later among the works it cites.
Penalising the biases in norm regularisation enforces sparsity
Etienne Boursier and Nicolas Flammarion · 2023
Later among the works it cites.
Learning a neuron by a shallow relu network: Dynamics and implicit bias for correlated inputs
Dmitry Chistikov, Matthias Englert, and Ranko Lazic · 2023
Later among the works it cites.
On the special role of class-selective neurons in early training
Omkar Ranadive, Nikhil Thakurdesai, Ari S. Morcos, Matthew L Leavitt, and Stephane Deny · 2023
Later among the works it cites.
Regression as classification: Influence of task formulation on neural network features
Lawrence Stewart, Francis Bach, Quentin Berthet, and Jean-Philippe Vert · 2023
Later among the works it cites.
Benign overfitting in ridge regression
Alexander Tsigler and Peter L Bartlett · 2023
Later among the works it cites.
On the spectral bias of two-layer linear networks
Aditya Vardhan Varre, Maria-Luiza Vladarean, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2023
Later among the works it cites.
Understanding multi-phase optimization dynamics and rich nonlinear behaviors of relu networks
Mingze Wang and Chao Ma · 2023
Later among the works it cites.
Simplicity bias and optimization threshold in two-layer relu networks
Etienne Boursier and Nicolas Flammarion · 2024
Closest in time.
SGD finds then tunes features in two-layer neural networks with near-optimal sample complexity: A case study in the XOR problem
Margalit Glasgow · 2024
Closest in time.
Directional convergence near small initializations and saddles in two-homogeneous neural networks
Akshay Kumar and Jarvis Haupt · 2024
Closest in time.
Early neuron alignment in two-layer reLU networks with small initialization
Hancheng Min, Enrique Mallada, and Rene Vidal · 2024
Closest in time.
Simplicity bias of two-layer networks beyond linearly separable data
Nikita Tsoy and Nikola Konstantinov · 2024
Closest in time.