Fetching the paper…
Reading the bibliography…
A key element of understanding the efficacy of overparameterized neural networks is characterizing how they represent functions as the number of weights in the network approaches infinity.
Depth-Width Tradeoffs in Approximating Natural Functions with Neural Networks
Itay Safran and Ohad Shamir · 1938
Earlier work this paper cites.
Asymptotic formulas for the dual Radon transform and applications
Donald C Solmon · 1987
Earlier work this paper cites.
Construction of neural nets using the Radon transform
Sean M. Carroll and Bradley W. Dickinson · 1989
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
Representation of functions by superpositions of a step or sigmoid function and their applications to neural network theory
Yoshifusa Ito · 1991
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R Barron · 1993
Earlier work this paper cites.
Approximation and estimation bounds for artificial neural networks
Andrew R Barron · 1994
Earlier work this paper cites.
For valid generalization the size of the weights is more important than the size of the network
Peter L Bartlett · 1997
Earlier work this paper cites.
Harmonic analysis of neural networks
Emmanuel J Candès · 1999
Earlier work this paper cites.
Ridgelets: A key to higher-dimensional intermittency?
Emmanuel J Candès and David L Donoho · 1999
Earlier work this paper cites.
The Radon transform
Sigurdur Helgason · 1999
Cited alongside, same era.
Approximation theory of the MLP model in neural networks
Allan Pinkus · 1999
Cited alongside, same era.
Convex neural networks
Yoshua Bengio, Nicolas L Roux, Pascal Vincent, Olivier Delalleau, and Patrice Marcotte · 2006
Cited alongside, same era.
Measure theory , volume 2
Vladimir I Bogachev · 2007
Cited alongside, same era.
Support theorems for the Radon transform and Cramér-Wold theorems
Jan Boman and Filip Lindskog · 2009
Cited alongside, same era.
Integration and probability , volume 157
Paul Malliavin · 2012
Cited alongside, same era.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Later among the works it cites.
Neural network with unbounded activation functions is universal approximator
Sho Sonoda and Noboru Murata · 2017
Later among the works it cites.
Error bounds for approximations with deep ReLU networks
Dmitry Yarotsky · 2017
Later among the works it cites.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Later among the works it cites.
Two models of double descent for weak features
Mikhail Belkin, Daniel Hsu, and Ji Xu · 2019
Closest in time.
Surprises in high-dimensional ridgeless least squares interpolation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Cited alongside, same era.
Norm-based capacity control in neural networks
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Cited alongside, same era.
Why Deep Neural Networks for Function Approximation?
Shiyu Liang and R. Srikant · 2016
Cited alongside, same era.
Benefits of depth in neural networks
Matus Telgarsky · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2019
Closest in time.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2019
Closest in time.
The generalization error of random features regression: Precise asymptotics and double descent curve
Song Mei and Andrea Montanari · 2019
Closest in time.
Lexicographic and depth-sensitive margins in homogeneous and non-homogeneous deep models
Mor Shpigel Nacson, Suriya Gunasekar, Jason D Lee, Nathan Srebro, and Daniel Soudry · 2019
Closest in time.
How do infinite width bounded norm networks look in function space?
Pedro Savarese, Itay Evron, Daniel Soudry, and Nathan Srebro · 2019
Closest in time.
Regularization matters: Generalization and optimization of neural nets v.s. their induced kernel
Colin Wei, Jason D Lee, Qiang Liu, and Tengyu Ma · 2019
Closest in time.