Fetching the paper…
Reading the bibliography…
We establish $ L^{\infty} $ and $ L^2 $ error bounds for functions of many variables that are approximated by linear combinations of ReLU (rectified linear unit) and squared ReLU ridge functions with $ \ell^1 $ and $ \ell^0 $ controls on their inner and outer parameters.
On the two different aspects of the representative method: The method of stratified sampling and the method of purposive selection
Jerzy Neyman · 1934
Earlier work this paper cites.
The uniform convergence of frequencies of the appearance of events to their probabilities
V. N. Vapnik and A. Ja. Červonenkis · 1971
Earlier work this paper cites.
Neural net approximation
Andrew R. Barron · 1992
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R. Barron · 1993
Earlier work this paper cites.
Hinging hyperplanes for regression, classification, and function approximation
Leo Breiman · 1993
Earlier work this paper cites.
Approximation and estimation bounds for artificial neural networks
Andrew R. Barron · 1994
Earlier work this paper cites.
Sphere packing numbers for subsets of the Boolean n n -cube with bounded Vapnik-Chervonenkis dimension
David Haussler · 1995
Earlier work this paper cites.
Sup-norm approximation bounds for networks through probabilistic methods
Joseph E. Yukich, Maxwell B. Stinchcombe, and Halbert White · 1995
Earlier work this paper cites.
Random approximants and neural networks
Y. Makovoz · 1996
Cited alongside, same era.
Weak convergence and empirical processes
Aad W. van der Vaart and Jon A. Wellner · 1996
Cited alongside, same era.
The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network
Peter L. Bartlett · 1998
Cited alongside, same era.
Uniform approximation by neural networks
Y. Makovoz · 1998
Cited alongside, same era.
A better approximation for balls
Gerald H. L. Cheang and Andrew R. Barron · 2000
Cited alongside, same era.
Estimates of covering numbers of convex sets with slowly decaying orthogonal subsets
Vĕra Ku̇rková and Marcello Sanguineti · 2007
Cited alongside, same era.
Concentration inequalities
Stéphane Boucheron, Gábor Lugosi, and Pascal Massart · 2013
Later among the works it cites.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
Majid Janzamin, Hanie Sedghi, and Anima Anandkumar · 2015
Later among the works it cites.
Learning halfspaces and neural networks with random initialization
Yuchen Zhang, Jason D Lee, Martin J Wainwright, and Michael I Jordan · 2015
Later among the works it cites.
Globally optimal gradient descent for a ConvNet with Gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Closest in time.
Learning combinations of sigmoids through gradient estimation
Stratis Ioannidis and Andrea Montanari · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Minimax rates of estimation for high-dimensional linear regression over ℓ q \ell_{q} -balls
Garvesh Raskutti, Martin J. Wainwright, and Bin Yu · 2011
Cited alongside, same era.
Minimax lower bounds for ridge combinations including neural nets
Jason M. Klusowski and Andrew R. Barron · 2017
Closest in time.
Risk bounds for high-dimensional ridge function combinations including neural networks
Jason M. Klusowski and Andrew R. Barron · 2018
Closest in time.