Fetching the paper…
Reading the bibliography…
While the universal approximation property holds both for hierarchical and shallow networks, we prove that deep (hierarchical) networks can approximate the class of compositional functions with the same accuracy as shallow networks but with exponentially lower number of training parameters as well as VC-dimension.
On direct and converse theorems in the theory of weighted polynomial approximation
Freud, G. (1972) · 1972
Earlier work this paper cites.
Neocognitron: A self-organizing neural network for a mechanism of pattern recognition unaffected by shift in position
Fukushima, K. (1980) · 1980
Earlier work this paper cites.
Computational Limitations for Small Depth Circuits
Hastad, J. T. (1987) · 1987
Earlier work this paper cites.
Optimal nonlinear approximation
DeVore, R. A., Howard, R., and Micchelli, C. A. (1989) · 1989
Earlier work this paper cites.
Constant depth circuits, fourier transform, and learnability
Linial, N., Y., M., and N., N. (1993) · 1993
Earlier work this paper cites.
Neural networks for localized approximation of real functions
Mhaskar, H. N. (1993) · 1993
Earlier work this paper cites.
Learning boolean functions via the fourier transform
Mansour, Y. (1994) · 1994
Earlier work this paper cites.
Neural networks for optimal approximation of smooth and analytic functions
Mhaskar, H. N. (1996) · 1996
Earlier work this paper cites.
Origins of scaling in natural images
Ruderman, D. (1997) · 1997
Earlier work this paper cites.
Representations properties of multilayer feedforward networks
B. Moore, B. and Poggio, T. (1998) · 1998
Earlier work this paper cites.
Approximation theory of the mlp model in neural networks
Pinkus, A. (1999) · 1999
Cited alongside, same era.
Neural Network Learning - Theoretical Foundations
Anthony, M. and Bartlett, P. (2002) · 2002
Cited alongside, same era.
On the degree of approximation in multivariate weighted approximation
Mhaskar, H. N. (2003) · 2003
Cited alongside, same era.
When is approximation by Gaussian networks necessarily a linear process?
Mhaskar, H. N. (2004) · 2004
Cited alongside, same era.
Scaling learning algorithms towards ai
Bengio, Y. and LeCun, Y. (2007) · 2007
Cited alongside, same era.
Shallow vs. deep sum-product networks
Delalleau, O. and Bengio, Y. (2011) · 2011
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2015) · 2015
Later among the works it cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C. (2015) · 2015
Later among the works it cites.
Deep learning
LeCun, Y., Bengio, Y., and G., H. (2015) · 2015
Later among the works it cites.
I-theory on depth vs width: hierarchical function composition
Poggio, T., Anselmi, F., and Rosasco, L. (2015) · 2015
Later among the works it cites.
Representation benefits of deep feedforward networks
Telgarsky, M. (2015) · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Livni, R., Shalev-Shwartz, S., and Shamir, O. (2013) · 2013
Cited alongside, same era.
On the number of linear regions of deep neural networks
Montufar, G. F.and Pascanu, R., Cho, K., and Bengio, Y. (2014) · 2014
Cited alongside, same era.
Unsupervised learning of invariant representations
Anselmi, F., Leibo, J. Z., Rosasco, L., Mutch, J., Tacchetti, A., and Poggio, T. (2015) · 2015
Cited alongside, same era.
Hierarchical models of object recognition in cortex
Riesenhuber, M. and Poggio, T. (1999a)
Cited in the paper.
Hierarchical models of object recognition in cortex
Riesenhuber, M. and Poggio, T. (1999b)
Cited in the paper.
Vedaldi, A. and Lenc, K. (2015) · 2015
Later among the works it cites.
Bridging the gap between residual learning, recurrent neural networks and visual cortex
Liao, Q. and Poggio, T. (2016) · 2016
Closest in time.
Learning real and boolean functions: When is deep better than shallow?
Mhaskar, H., Liao, Q., and Poggio, T. (2016) · 2016
Closest in time.
Soatto, S. (2011) · 2053
Closest in time.