Fetching the paper…
Reading the bibliography…
The analysis of neural network training beyond their linearization regime remains an outstanding open question, even in the simplest setup of a single hidden-layer.
Two models of double descent for weak features
Mikhail Belkin, Daniel Hsu, and Ji Xu · 1903
Earlier work this paper cites.
Implicit regularization for deep neural networks driven by an Ornstein-Uhlenbeck like process
Guy Blanc, Neha Gupta, Gregory Valiant, and Paul Valiant · 1904
Earlier work this paper cites.
A Function Space View of Bounded Norm Infinite Width ReLU Nets: The Multivariate Case
Greg Ongie, Rebecca Willett, Daniel Soudry, and Nathan Srebro · 1910
Earlier work this paper cites.
Remarks on problems in approximation theory
S. Zuhovickii · 1948
Earlier work this paper cites.
Spline solutions to l1 extremal problems in one and several variables
SD Fisher and Joseph W Jerome · 1975
Earlier work this paper cites.
For valid generalization the size of the weights is more important than the size of the network
Peter L Bartlett · 1997
Earlier work this paper cites.
Dynamics of training
Siegfried Bös and Manfred Opper · 1997
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Norm-Based Capacity Control in Neural Networks
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Earlier work this paper cites.
High-dimensional dynamics of generalization error in neural networks
Madhu S Advani and Andrew M Saxe · 2017
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Earlier work this paper cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Cited alongside, same era.
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2017
Cited alongside, same era.
A note on lazy training in supervised differentiable programming
Lenaic Chizat and Francis Bach · 2018
Cited alongside, same era.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lenaic Chizat and Francis Bach · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
On representer theorems and convex regularization
Claire Boyer, Antonin Chambolle, Yohann De Castro, Vincent Duval, Frédéric De Gournay, and Pierre Weiss · 2019
Later among the works it cites.
Reconciling modern machine learning practice and the bias-variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Later among the works it cites.
Sparse optimization on measures with over-parameterized gradient descent
Lenaic Chizat · 2019
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
Song Mei and Andrea Montanari · 2019
Later among the works it cites.
Barron spaces and the compositional function spaces for neural network models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gradient Descent Quantizes ReLU Network Features
Hartmut Maennel, Olivier Bousquet, and Sylvain Gelly · 2018
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Cited alongside, same era.
A modern take on the bias-variance tradeoff in neural networks
Brady Neal, Sarthak Mittal, Aristide Baratin, Vinayak Tantia, Matthew Scicluna, Simon Lacoste-Julien, and Ioannis Mitliagkas · 2018
Cited alongside, same era.
Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks
Grant Rotskoff and Eric Vanden-Eijnden · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Cited alongside, same era.
Dgm: A deep learning algorithm for solving partial differential equations
Justin Sirignano and Konstantinos Spiliopoulos · 2018
Cited alongside, same era.
Chao Ma, Lei Wu, and Weinan E · 2019
Later among the works it cites.
How do infinite width bounded norm networks look in function space?
Pedro Savarese, Itay Evron, Daniel Soudry, and Nathan Srebro · 2019
Later among the works it cites.
Gradient dynamics of shallow univariate relu networks
Francis Williams, Matthew Trager, Daniele Panozzo, Claudio Silva, Denis Zorin, and Joan Bruna · 2019
Later among the works it cites.
Network size and weights size for memorization with two-layers neural networks
Sébastien Bubeck, Ronen Eldan, Yin Tat Lee, and Dan Mikulincer · 2020
Closest in time.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lénaïc Chizat and Francis Bach · 2020
Closest in time.
Convex geometry and duality of over-parameterized neural networks
Tolga Ergen and Mert Pilanci · 2020
Closest in time.
Superpolynomial lower bounds for learning one-layer neural networks using gradient descent
S. Goel, A. Gollakota, Z. Jin, S. Karmalkar, and A. Klivans · 2020
Closest in time.