Fetching the paper…
Reading the bibliography…
The paper briefy reviews several recent results on hierarchical architectures for learning from examples, that may formally explain the conditions under which Deep Convolutional Neural Networks perform much better in function approximation problems than shallow, one-hidden layer architectures.
On direct and converse theorems in the theory of weighted polynomial approximation
G. Freud · 1972
Earlier work this paper cites.
Orthogonal polynomials
G. Szegö · 1975
Earlier work this paper cites.
Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position
K. Fukushima · 1980
Earlier work this paper cites.
Computational Limitations for Small Depth Circuits
J. T. Hastad · 1987
Earlier work this paper cites.
Optimal nonlinear approximation
R. A. DeVore, R. Howard, and C. A. Micchelli · 1989
Earlier work this paper cites.
An introduction to wavelets
C. K. Chui · 1992
Earlier work this paper cites.
Neural networks for localized approximation
C. K. Chui, X. Li, and H. N. Mhaskar · 1994
Earlier work this paper cites.
Limitations of the approximation capabilities of neural networks with one hidden layer
C. K. Chui, X. Li, and H. N. Mhaskar · 1996
Earlier work this paper cites.
Neural networks for optimal approximation of smooth and analytic functions
H. N. Mhaskar · 1996
Earlier work this paper cites.
Origins of scaling in natural images
D. Ruderman · 1997
Cited alongside, same era.
Special functions
G. E. Andrews, R. Askey, and R. Roy · 1999
Cited alongside, same era.
Hierarchical models of object recognition in cortex
M. Riesenhuber and T. Poggio · 1999
Cited alongside, same era.
High-dimensional data analysis: The curses and blessings of dimensionality
D. L. Donoho et al · 2000
Cited alongside, same era.
On the degree of approximation in multivariate weighted approximation
H. N. Mhaskar · 2003
Cited alongside, same era.
The mathematics of learning: Dealing with data
T. Poggio and S. Smale · 2003
Cited alongside, same era.
When is approximation by Gaussian networks necessarily a linear process?
Eignets for function approximation on manifolds
H. N. Mhaskar · 2010
Later among the works it cites.
The power of depth for feedforward neural networks
R. Eldan and O. Shamir · 2015
Later among the works it cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Later among the works it cites.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Later among the works it cites.
I-theory on depth vs width: hierarchical function composition
T. Poggio, F. Anselmi, and L. Rosasco · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. N. Mhaskar · 2004
Cited alongside, same era.
A Markov-Bernstein inequality for Gaussian networks
H. N. Mhaskar · 2005
Cited alongside, same era.
Weighted quadrature formulas and approximation by zonal function networks on the sphere
H. N. Mhaskar · 2006
Cited alongside, same era.
Local approximation using Hermite functions
H. N. Mhaskar
Cited in the paper.
M. Telgarsky · 2015
Later among the works it cites.
Learning real and boolean functions: When is deep better than shallow
H. N. Mhaskar, Q. Liao, and T. Poggio · 2016
Closest in time.
Singular integrals and differentiability properties of functions (PMS-30)
E. M. Stein · 2016
Closest in time.