Fetching the paper…
Reading the bibliography…
We theoretically discuss why deep neural networks (DNNs) performs better than other models in some cases by investigating statistical properties of DNNs for non-smooth functions.
Metric entropy of some classes of sets with differentiable boundaries
Richard M Dudley · 1974
Earlier work this paper cites.
Optimal global rates of convergence for nonparametric regression
CJ Stone · 1982
Earlier work this paper cites.
Mean integrated squared error of kernel estimators when the density and its derivative are not necessarily continuous
Constance van Eeden · 1985
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
Pierre Baldi and Kurt Hornik · 1989
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R Barron · 1993
Earlier work this paper cites.
Kernel-type estimators of jump points and values of a regression function
JS Wu and CK Chu · 1993
Earlier work this paper cites.
Nonparametric function estimation and bandwidth selection for discontinuous regression functions
JS Wu and CK Chu · 1993
Earlier work this paper cites.
Approximation and estimation bounds for artificial neural networks
Andrew R Barron · 1994
Earlier work this paper cites.
Asymptotical minimax recovery of sets with smooth boundaries
E Mammen and AB Tsybakov · 1995
Earlier work this paper cites.
Weak Convergence and Empirical Processes: With Applications to Statistics
AW van der Vaart and Jon Wellner · 1996
Earlier work this paper cites.
The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network
Peter L Bartlett · 1998
Earlier work this paper cites.
Smooth discrimination analysis
Enno Mammen, Alexandre B Tsybakov, et al · 1999
Earlier work this paper cites.
Information-theoretic determination of minimax rates of convergence
Yuhong Yang and Andrew Barron · 1999
Earlier work this paper cites.
Local minima and plateaus in hierarchical structures of multilayer perceptrons
Kenji Fukumizu and Shun-ichi Amari · 2000
Earlier work this paper cites.
Recovering edges in ill-posed inverse problems: Optimality of curvelet frames
Emmanuel J Candes and David L Donoho · 2002
Earlier work this paper cites.
New tight frames of curvelets and optimal representations of objects with piecewise c2 singularities
Emmanuel J Candès and David L Donoho · 2004
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Geoffrey E Hinton, Simon Osindero, and Yee-Whye Teh · 2006
Earlier work this paper cites.
Local rademacher complexities and oracle inequalities in risk minimization
Vladimir Koltchinskii · 2006
Cited alongside, same era.
All of nonparametric statistics: with 52 illustrations
Larry Alan Wasserman · 2006
Cited alongside, same era.
A unified architecture for natural language processing: Deep neural networks with multitask learning
Ronan Collobert and Jason Weston · 2008
Cited alongside, same era.
Support vector machines
Ingo Steinwart and Andreas Christmann · 2008
Cited alongside, same era.
Rates of contraction of posterior distributions based on gaussian process priors
AW van der Vaart and JH van Zanten · 2008
Cited alongside, same era.
Neural network learning: Theoretical foundations
Martin Anthony and Peter L Bartlett · 2009
Cited alongside, same era.
On the number of linear regions of deep neural networks
Guido F Montufar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio · 2014
Later among the works it cites.
Mathematical foundations of infinite-dimensional statistical models
Evarist Giné and Richard Nickl · 2015
Later among the works it cites.
Probabilistic backpropagation for scalable learning of bayesian neural networks
José Miguel Hernández-Lobato and Ryan Adams · 2015
Later among the works it cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Later among the works it cites.
Norm-based capacity control in neural networks
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Later among the works it cites.
Deep learning in neural networks: An overview
Jürgen Schmidhuber · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Introduction to nonparametric estimation, 2009
Alexandre B Tsybakov · 2009
Cited alongside, same era.
On the expressive power of deep architectures
Yoshua Bengio and Olivier Delalleau · 2011
Cited alongside, same era.
Compactly supported shearlets are optimally sparse
Gitta Kutyniok and Wang-Q Lim · 2011
Cited alongside, same era.
On optimization methods for deep learning
Quoc V Le, Jiquan Ngiam, Adam Coates, Abhik Lahiri, Bobby Prochnow, and Andrew Y Ng · 2011
Cited alongside, same era.
Information rates of nonparametric gaussian process methods
Aad van der Vaart and Harry van Zanten · 2011
Cited alongside, same era.
Stochastic expansions using continuous dictionaries: Lévy adaptive regression kernels
Robert L Wolpert, Merlise A Clyde, and Chong Tu · 2011
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Later among the works it cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Later among the works it cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Daniel Soudry and Yair Carmon · 2016
Later among the works it cites.
Memory-optimal neural network approximation
Helmut Bölcskei, Philipp Grohs, Gitta Kutyniok, and Philipp Petersen · 2017
Later among the works it cites.
Optimal approximation with sparsely connected deep neural networks
Helmut Bölcskei, Philipp Grohs, Gitta Kutyniok, and Philipp Petersen · 2017
Later among the works it cites.
Optimal approximation of piecewise smooth functions using deep relu neural networks
Philipp Petersen and Felix Voigtlaender · 2017
Later among the works it cites.
Nonparametric regression using deep neural networks with relu activation function
Johannes Schmidt-Hieber · 2017
Later among the works it cites.
Error bounds for approximations with deep relu networks
Dmitry Yarotsky · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Later among the works it cites.
Statistically efficient estimation for non-smooth probability densities
Masaaki Imaizumi, Takanori Maehara, and Yuichi Yoshida · 2018
Closest in time.
Fast generalization error bound of deep learning from a kernel perspective
Taiji Suzuki · 2018
Closest in time.