Fetching the paper…
Reading the bibliography…
The paper characterizes classes of functions for which deep learning can be exponentially better than shallow learning.
E. Corominas and F. S. Balaguer, “Condiciones para que una funcion infinitamente derivable sea un polinomio,”
1954
Earlier work this paper cites.
Cambridge MA: The MIT Press, ISBN 0-262-63022-2, 1972
M. Minsky and S. Papert, · 1972
Earlier work this paper cites.
K. Fukushima, “Neocognitron: A self-organizing neural network for a mechanism of pattern recognition unaffected by shift in position,”
1980
Earlier work this paper cites.
T. Poggio and W. Reichardt, “On the representation of multi-input systems: Computational properties of polynomial algorithms.,”
1980
Earlier work this paper cites.
MIT Press, 1987
J. T. Hastad, · 1987
Earlier work this paper cites.
R. A. DeVore, R. Howard, and C. A. Micchelli, “Optimal nonlinear approximation,”
1989
Earlier work this paper cites.
T. Poggio and F. Girosi, “A theory of networks for approximation and learning,”
1989
Earlier work this paper cites.
F. Girosi and T. Poggio, “Representation properties of networks: Kolmogorov’s theorem is irrelevant,”
1989
Earlier work this paper cites.
F. Girosi and T. Poggio, “Networks and the best approximation property,”
1990
Earlier work this paper cites.
J. Mihalik, “Hierarchical vector quantization. of images in transform domain.,”
1992
Earlier work this paper cites.
H. Mhaskar, “Approximation properties of a multilayered feedforward artificial neural network,”
1993
Earlier work this paper cites.
H. N. Mhaskar, “Neural networks for localized approximation of real functions,” in
1993
Earlier work this paper cites.
N. Linial, M. Y., and N. N., “Constant depth circuits, fourier transform, and learnability,”
1993
Earlier work this paper cites.
C. Chui, X. Li, and H. Mhaskar, “Neural networks for localized approximation,”
1994
Earlier work this paper cites.
Y. Mansour, “Learning boolean functions via the fourier transform,” in
1994
Earlier work this paper cites.
F. Girosi, M. Jones, and T. Poggio, “Regularization theory and neural networks architectures,”
1995
Earlier work this paper cites.
C. K. Chui, X. Li, and H. N. Mhaskar, “Limitations of the approximation capabilities of neural networks with one hidden layer,”
1996
Earlier work this paper cites.
H. N. Mhaskar, “Neural networks for optimal approximation of smooth and analytic functions,”
1996
Earlier work this paper cites.
D. Ruderman, “Origins of scaling in natural images,”
1997
Earlier work this paper cites.
B. B. Moore and T. Poggio, “Representations properties of multilayer feedforward networks,”
1998
Cited alongside, same era.
M. Riesenhuber and T. Poggio, “Hierarchical models of object recognition in cortex,”
1999
Cited alongside, same era.
A. Pinkus, “Approximation theory of the mlp model in neural networks,”
1999
Cited alongside, same era.
D. L. Donoho, “High-dimensional data analysis: The curses and blessings of dimensionality,” in
2000
Cited alongside, same era.
Cambridge University Press, 2002
M. Anthony and P. Bartlett, · 2002
Cited alongside, same era.
T. Poggio and S. Smale, “The mathematics of learning: Dealing with data,”
2003
Cited alongside, same era.
T. Poggio, L. Rosasco, A. Shashua, N. Cohen, and F. Anselmi, “Notes on hierarchical splines, dclns and i-theory,” tech. rep., MIT Computer Science and Artificial Intelligence Laboratory, 2015
2015
Later among the works it cites.
T. Poggio, F. Anselmi, and L. Rosasco, “I-theory on depth vs width: hierarchical function composition,”
2015
Later among the works it cites.
Y. LeCun, Y. Bengio, and H. G., “Deep learning,”
2015
Later among the works it cites.
N. Cohen, O. Sharir, and A. Shashua, “On the expressive power of deep learning: a tensor analysis,”
2015
Later among the works it cites.
F. Anselmi, J. Z. Leibo, L. Rosasco, J. Mutch, A. Tacchetti, and T. Poggio, “Unsupervised learning of invariant representations,”
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. N. Mhaskar, “On the tractability of multivariate integration and approximation by neural networks,”
2004
Cited alongside, same era.
Y. Bengio and Y. LeCun, “Scaling learning algorithms towards ai,” in
2007
Cited alongside, same era.
L. Grasedyck, “Hierarchical Singular Value Decomposition of Tensors,”
2010
Cited alongside, same era.
O. Delalleau and Y. Bengio, “Shallow vs. deep sum-product networks,” in
2011
Cited alongside, same era.
2011
Cited alongside, same era.
J. Bergstra and Y. Bengio, “Random search for hyper-parameter optimization,”
2012
Cited alongside, same era.
T. Poggio, L. Rosaco, A. Shashua, N. Cohen, and F. Anselmi, “Notes on hierarchical splines, dclns and i-theory,”
2015
Later among the works it cites.
M. Telgarsky, “Representation benefits of deep feedforward networks,”
2015
Later among the works it cites.
Software available from tensorflow.org
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng, “TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015 · 2015
Later among the works it cites.
B. M. Lake, R. Salakhutdinov, and J. B. Tenenabum, “Human-level concept learning through probabilistic program induction,”
2015
Later among the works it cites.
A. Maurer, “Bounds for Linear Multi-Task Learning,”
2015
Later among the works it cites.
F. Anselmi, L. Rosasco, and T. Tan, C.and Poggio, “Deep Convolutional Networks are Hierarchical Kernel Machines,”
2015
Later among the works it cites.
H. Mhaskar, Q. Liao, and T. Poggio, “Learning real and boolean functions: When is deep better than shallow?,”
2016
Closest in time.
H. Mhaskar and T. Poggio, “Deep versus shallow networks: an approximation theory perspective,”
2016
Closest in time.
Q. Liao and T. Poggio, “Bridging the gap between residual learning, recurrent neural networks and visual cortex,”
2016
Closest in time.
2016
Closest in time.
R. Eldan and O. Shamir, “The power of depth for feedforward neural networks,”
2016
Closest in time.
M. Lin, H.and Tegmark, “Why does deep and cheap learning work so well?,”
2016
Closest in time.
MIT Press, 2016
F. Anselmi and T. Poggio, · 2016
Closest in time.