Fetching the paper…
Reading the bibliography…
We consider neural networks with a single hidden layer and non-decreasing homogeneous activa-tion functions like the rectified linear units.
Analytic extensions of differentiable functions defined in closed sets
H. Whitney · 1934
Earlier work this paper cites.
An algorithm for quadratic programming
M. Frank and P. Wolfe · 1956
Earlier work this paper cites.
The perceptron: a probabilistic model for information storage and organization in the brain
F. Rosenblatt · 1958
Earlier work this paper cites.
On the stationary values of a second-degree polynomial on the unit sphere
G. E. Forsythe and G. H. Golub · 1965
Earlier work this paper cites.
The minimization of a smooth convex functional on a convex set
V. F. Dem’yanov and A. M. Rubinov · 1967
Earlier work this paper cites.
Zu einem problem von shephard über die projektionen konvexer körper
R. Schneider · 1967
Earlier work this paper cites.
A class of convex bodies
E. D. Bolker · 1969
Earlier work this paper cites.
Conditional gradient algorithms with open loop step size rules
J. C. Dunn and S. Harshbarger · 1978
Earlier work this paper cites.
Projection pursuit regression
J. H. Friedman and W. Stuetzle · 1981
Earlier work this paper cites.
Learning representations by back-propagating errors
D. E. Rumelhart, G. E Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Algorithms in Combinatorial Geometry , volume 10
H. Edelsbrunner · 1987
Earlier work this paper cites.
Real and Complex Analysis
W. Rudin · 1987
Earlier work this paper cites.
Projection bodies
J. Bourgain and J. Lindenstrauss · 1988
Earlier work this paper cites.
Approximation of zonoids by zonotopes
J. Bourgain, J. Lindenstrauss, and V. Milman · 1989
Earlier work this paper cites.
Optimal nonlinear approximation
R. A. DeVore, R. Howard, and C. Micchelli · 1989
Earlier work this paper cites.
Generalized Additive Models
T. J. Hastie and R. J. Tibshirani · 1990
Earlier work this paper cites.
Measure Theory and Fine Properties of Functions , volume 5
L. C. Evans and R. F. Gariepy · 1991
Earlier work this paper cites.
Sliced inverse regression for dimension reduction
K.-C. Li · 1991
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
A. R. Barron · 1993
Earlier work this paper cites.
Hinging hyperplanes for regression, classification, and function approximation
L. Breiman · 1993
Earlier work this paper cites.
Multilayer feedforward networks with a nonpolynomial activation function can approximate any function
M. Leshno, V. Y. Lin, A. Pinkus, and S. Schocken · 1993
Earlier work this paper cites.
Neural Networks: A Comprehensive Foundation
S. Haykin · 1994
Earlier work this paper cites.
Bayesian Learning for Neural Networks
R. M. Neal · 1995
Earlier work this paper cites.
Efficient agnostic learning of neural networks with bounded fan-in
W. S. Lee, P. L. Bartlett, and R. C. Williamson · 1996
Earlier work this paper cites.
Improved upper bounds for approximation by zonotopes
J. Matoušek · 1996
Earlier work this paper cites.
Generative models for discovering sparse distributed representations
G. E. Hinton and Z. Ghahramani · 1997
Earlier work this paper cites.
Convex Analysis
R. T. Rockafellar · 1997
Earlier work this paper cites.
Uniform approximation by neural networks
Y. Makovoz · 1998
Earlier work this paper cites.
Semidefinite relaxation and nonconvex quadratic optimization
Y. Nesterov · 1998
Earlier work this paper cites.
Approximation by ridge functions and neural networks
P. P. Petrushev · 1998
Cited alongside, same era.
Approximation theory of the MLP model in neural networks
A. Pinkus · 1999
Cited alongside, same era.
On the near optimality of the stochastic approximation of smooth functions by neural networks
V. E. Maiorov and R. Meir · 2000
Cited alongside, same era.
Error bounds for approximation with neural networks
M. Burger and A. Neubauer · 2001
Cited alongside, same era.
Rademacher penalties and structural risk minimization
V. Koltchinskii · 2001
Cited alongside, same era.
Bounds on rates of variable-basis and neural-network approximation
V. Kurkova and M. Sanguineti · 2001
Cited alongside, same era.
Random features for large-scale kernel machines
A. Rahimi and B. Recht · 2007
Later among the works it cites.
ℓ 1 \ell_{1} -regularization in infinite dimensional feature spaces
S. Rosset, G. Swirszcz, N. Srebro, and J. Zhu · 2007
Later among the works it cites.
The tradeoffs of large scale learning
L. Bottou and O. Bousquet · 2008
Later among the works it cites.
A new algorithm for estimating the effective dimension-reduction subspace
A. S. Dalalyan, A. Juditsky, and V. Spokoiny · 2008
Later among the works it cites.
SpAM: Sparse additive models
P. Ravikumar, H. Liu, J. Lafferty, and L. Wasserman · 2008
Later among the works it cites.
Optimization Algorithms on Matrix Manifolds
P.-A. Absil, R. Mahony, and R. Sepulchre · 2009
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Regularization with dot-product kernels
A. J. Smola, Z. L. Ovari, and R. C. Williamson · 2001
Cited alongside, same era.
A Course in Convexity , volume 54
A. Barvinok · 2002
Cited alongside, same era.
A Distribution-free Theory of Nonparametric Regression
L. Györfi and A. Krzyzak · 2002
Cited alongside, same era.
Sobolev Spaces , volume 140
R. A. Adams and J. F. Fournier · 2003
Cited alongside, same era.
Rademacher and Gaussian complexities: Risk bounds and structural results
P. L. Bartlett and S. Mendelson · 2003
Cited alongside, same era.
Zonotopes as bounding volumes
L. J. Guibas, A. Nguyen, and L. Zhang · 2003
Cited alongside, same era.
Kernel methods for deep learning
Y. Cho and L. K. Saul · 2009
Later among the works it cites.
Hardness of learning halfspaces with noise
V. Guruswami and P. Raghavendra · 2009
Later among the works it cites.
The Elements of Statistical Learning
T. Hastie, R. Tibshirani, and J. Friedman · 2009
Later among the works it cites.
On the complexity of linear prediction: Risk bounds, margin bounds, and regularization
S. M. Kakade, K. Sridharan, and A. Tewari · 2009
Later among the works it cites.
Rectified linear units improve restricted boltzmann machines
V. Nair and G. E. Hinton · 2010
Later among the works it cites.
Statistics for high-dimensional data: methods, theory and applications
P. Bühlmann and S. Van De Geer · 2011
Later among the works it cites.
Spherical Harmonics and Approximations on the Unit Sphere: an Introduction , volume 2044
K. Atkinson and W. Han · 2012
Later among the works it cites.
Lifted coordinate descent for learning with trace-norm regularization
M. Dudik, Z. Harchaoui, and J. Malick · 2012
Later among the works it cites.
Spherical Harmonics in p p Dimensions
C. Frye and C. J. Efthimiou · 2012
Later among the works it cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Later among the works it cites.
Learning from an Optimization Viewpoint
K. Sridharan · 2012
Later among the works it cites.
Accelerated training for matrix-norm regularization: A boosting approach
X. Zhang, D. Schuurmans, and Y. Yu · 2012
Later among the works it cites.
Convex relaxations of structured matrix factorizations
F. Bach · 2013
Later among the works it cites.
Smoothing Spline ANOVA Models , volume 297
C. Gu · 2013
Later among the works it cites.
Conditional gradient algorithms for norm-regularized smooth convex optimization
Z. Harchaoui, A. Juditsky, and A. Nemirovski · 2013
Later among the works it cites.
Revisiting Frank-Wolfe: Projection-free sparse convex optimization
M. Jaggi · 2013
Later among the works it cites.
The complexity of large-scale convex programming under a linear optimization oracle
G. Lan · 2013
Later among the works it cites.
Duality between subgradient and conditional gradient methods
F. Bach · 2014
Closest in time.
Computational aspects of the Hausdorff distance in unbounded dimension
S. König · 2014
Closest in time.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Closest in time.
On the number of linear regions of deep neural networks
G. F. Montufar, R. Pascanu, K. Cho, and Y. Bengio · 2014
Closest in time.
Understanding Machine Learning: From Theory to Algorithms
S. Shalev-Shwartz and S. Ben-David · 2014
Closest in time.
On the equivalence between quadrature rules and random features
F. Bach · 2015
Closest in time.