Fetching the paper…
Reading the bibliography…
It has been experimentally observed in recent years that multi-layer artificial neural networks have a surprising ability to generalize, even when trained with far more parameters than observations.
Geometrical and Statistical Properties of Linear Threshold Devices
Thomas M. Cover · 1964
Earlier work this paper cites.
Pattern recognition
Thomas M. Cover · 1968
Earlier work this paper cites.
An introduction to probability theory and its applications
William Feller · 1971
Earlier work this paper cites.
The theory of error-correcting codes
Florence Jessie MacWilliams and Neil James Alexander Sloane · 1977
Earlier work this paper cites.
Lower bounds for constant weight codes
R. L. Graham and N. J. A. Sloane · 1980
Earlier work this paper cites.
Statistical learning networks: A unifying view
Andrew R Barron and Roger L Barron · 1988
Earlier work this paper cites.
On the capabilities of multilayer perceptrons
Eric B Baum · 1988
Earlier work this paper cites.
Approximation and estimation bounds for artificial neural networks
Andrew R Barron · 1991
Earlier work this paper cites.
Neural net approximation
Andrew R Barron · 1992
Earlier work this paper cites.
Sets of matrices all infinite products of which converge
Ingrid Daubechies and Jeffrey C Lagarias · 1992
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R Barron · 1993
Earlier work this paper cites.
Hinging hyperplanes for regression, classification, and function approximation
Leo Breiman · 1993
Earlier work this paper cites.
Approximation and estimation bounds for artificial neural networks
Andrew R Barron · 1994
Earlier work this paper cites.
Neural network approximation and estimation of functions
Gerald L.H. Cheang · 1998
Earlier work this paper cites.
Introduction to large truncated Toeplitz matrices
Albrecht Böttcher and Bernd Silbermann · 1999
Earlier work this paper cites.
Estimation with two hidden layer neural nets
Gerald L Cheang and Andrew R Barron · 1999
Earlier work this paper cites.
Information-theoretic determination of minimax rates of convergence
Yuhong Yang and Andrew R Barron · 1999
Cited alongside, same era.
A local Riemann hypothesis, I
Daniel Bump, Kwok-Kwong Choi, Pär Kurlberg, and Jeffrey Vaaler · 2000
Cited alongside, same era.
On the size of convex hulls of small sets
Shahar Mendelson · 2002
Cited alongside, same era.
Introduction to nonparametric estimation
Alexandre B. Tsybakov · 2004
Cited alongside, same era.
Approximation and learning by greedy algorithms
Andrew R Barron, Albert Cohen, Wolfgang Dahmen, and Ronald A. DeVore · 2008
Cited alongside, same era.
Risk of penalized least squares, greedy selection and ℓ 1 \ell_{1} penalization for flexible function libraries
Cong Huang, G. LH Cheang, and Andrew R Barron · 2008
Cited alongside, same era.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Later among the works it cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Later among the works it cites.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Later among the works it cites.
Size-independent sample complexity of neural networks
Noah Golowich, Alexander Rakhlin, and Ohad Shamir · 2017
Later among the works it cites.
Nearly-tight VC-dimension bounds for piecewise linear neural networks
Nick Harvey, Christopher Liaw, and Abbas Mehrabian · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Minimax rates of estimation for high-dimensional linear regression over ℓ q \ell_{q} -balls
Garvesh Raskutti, Martin J. Wainwright, and Bin Yu · 2011
Cited alongside, same era.
Exponential screening and optimal rates of sparse estimation
Philippe Rigollet and Alexandre B. Tsybakov · 2011
Cited alongside, same era.
The convex geometry of linear inverse problems
Venkat Chandrasekaran, Benjamin Recht, Pablo A Parrilo, and Alan S Willsky · 2012
Cited alongside, same era.
Matrix analysis
Roger A. Horn and Charles R. Johnson · 2012
Cited alongside, same era.
Foundations of Machine Learning
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar · 2012
Cited alongside, same era.
Metric entropy and sparse linear approximation of ℓ q \ell_{q} -hulls for 0 < q ≤ 1 0<q\leq 1
Fuchang Gao, Ching-Kang Ing, and Yuhong Yang · 2013
Cited alongside, same era.
Kenji Kawaguchi, Leslie Pack Kaelbling, and Yoshua Bengio · 2017
Later among the works it cites.
Minimax lower bounds for ridge combinations including neural nets
Jason M Klusowski and Andrew R Barron · 2017
Later among the works it cites.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nati Srebro · 2017
Later among the works it cites.
Geometry of optimization and implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, Ruslan Salakhutdinov, and Nathan Srebro · 2017
Later among the works it cites.
Nonparametric regression using deep neural networks with ReLU activation function
Johannes Schmidt-Hieber · 2017
Later among the works it cites.
Error bounds for approximations with deep ReLU networks
Dmitry Yarotsky · 2017
Later among the works it cites.
Stronger generalization bounds for deep nets via a compression approach
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang · 2018
Closest in time.
Approximation by combinations of ReLU and squared ReLU ridge functions with ℓ 1 \ell^{1} and ℓ 0 \ell^{0} controls
Jason M Klusowski and Andrew R Barron · 2018
Closest in time.
Risk bounds for high-dimensional ridge function combinations including neural networks
Jason M Klusowski and Andrew R Barron · 2018
Closest in time.
Optimal approximation of continuous functions by very deep ReLU networks
Dmitry Yarotsky · 2018
Closest in time.