Fetching the paper…
Reading the bibliography…
This paper develops fundamental limits of deep neural network learning by characterizing what is possible if no constraints are imposed on the learning algorithm and on the amount of training data.
1905
Earlier work this paper cites.
1906
Earlier work this paper cites.
1910
Earlier work this paper cites.
W. McCulloch and W. Pitts, “A logical calculus of ideas immanent in nervous activity,” Bull. Math. Biophys. , vol. 5, pp. 115–133, 1943
1943
Earlier work this paper cites.
M. H. Stone, “The generalized Weierstrass approximation theorem,” Mathematics Magazine , vol. 21, pp. 167–184, 1948
1948
Earlier work this paper cites.
A. N. Kolmogorov, “On the representation of continuous functions of many variables by superposition of continuous functions of one variable and addition,” Dokl. Akad. Nauk SSSR , vol. 114, no. 5, pp. 953–956, 1957
1957
Earlier work this paper cites.
A. Kolmogorov and V. Tikhomirov, “ ε \varepsilon -entropy and ε \varepsilon -capacity of sets in function spaces,” Uspekhi Mat. Nauk. , vol. 14, no. 2, pp. 3–86, 1959
1959
Earlier work this paper cites.
R. T. Prosser, “The ε \varepsilon -entropy and ε \varepsilon -capacity of certain time-varying channels,” Journal of Mathematical Analysis and Applications , vol. 16, pp. 553–573, 1966
1966
Earlier work this paper cites.
H. G. Feichtinger, “On a new Segal algebra,” Monatshefte für Mathematik , vol. 92, pp. 269–289, 1981
1981
Earlier work this paper cites.
C. L. Fefferman, “The uncertainty principle,” Bull. Amer. Math. Soc. (N.S.) , vol. 9, no. 2, pp. 129–206, 1983. [Online]. Available: https://doi.org/10.1090/S0273-0979-1983-15154-6
1983
Earlier work this paper cites.
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning representations by back-propagating errors,” Nature , vol. 323, no. 6088, pp. 533–536, Oct. 1986. [Online]. Available: http://dx.doi.org/10.1038/323533a0
1986
Earlier work this paper cites.
G. Cybenko, “Approximation by superpositions of a sigmoidal function,” Mathematics of Control, Signals, and Systems , vol. 2, no. 4, pp. 303–314, 1989. [Online]. Available: http://dx.doi.org/10.1007/BF02551274
1989
Earlier work this paper cites.
K. Hornik, M. Stinchcombe, and H. White, “Multilayer feedforward networks are universal approximators,” Neural Networks , vol. 2, no. 5, pp. 359–366, 1989
1989
Earlier work this paper cites.
K.-I. Funahashi, “On the approximate realization of continuous mappings by neural networks,” Neural Networks , vol. 2, no. 3, pp. 183–192, 1989. [Online]. Available: //www.sciencedirect.com/science/article/pii/0893608089900038
1989
Earlier work this paper cites.
S. Mallat, “Multiresolution approximations and wavelet orthonormal bases of L 2 ( R ) L^{2}(R) ,” Trans. Amer. Math. Soc. , vol. 315, no. 1, pp. 69–87, Sep. 1989
1989
Earlier work this paper cites.
K. Hornik, “Approximation capabilities of multilayer feedforward networks,” Neural Networks , vol. 4, no. 2, pp. 251 – 257, 1991. [Online]. Available: http://www.sciencedirect.com/science/article/pii/089360809190009T
1991
Earlier work this paper cites.
I. Daubechies, Ten Lectures on Wavelets . SIAM, 1992
1992
Earlier work this paper cites.
C. K. Chui and J.-Z. Wang, “On compactly supported spline wavelets and a duality principle,” Trans. Amer. Math. Soc. , vol. 330, no. 2, pp. 903–915, Apr. 1992
1992
Earlier work this paper cites.
D. L. Donoho, “Unconditional bases are optimal bases for data compression and for statistical estimation,” Appl. Comput. Harmon. Anal. , vol. 1, no. 1, pp. 100 – 115, 1993. [Online]. Available: http://www.sciencedirect.com/science/article/pii/S1063520383710080
1993
Earlier work this paper cites.
A. R. Barron, “Universal approximation bounds for superpositions of a sigmoidal function,” IEEE Transactions on Information Theory , vol. 39, no. 3, pp. 930–945, 1993
1993
Earlier work this paper cites.
H. N. Mhaskar, “Approximation properties of a multilayered feedforward artificial neural network,” Advances in Computational Mathematics , vol. 1, no. 1, pp. 61–80, Feb 1993. [Online]. Available: https://doi.org/10.1007/BF02070821
1993
Earlier work this paper cites.
R. A. DeVore and G. G. Lorentz, Constructive Approximation . Springer, 1993
1993
Earlier work this paper cites.
C. L. Fefferman, “Reconstructing a neural net from its output,” Revista Matemática Iberoamericana , vol. 10, no. 3, pp. 507–555, 1994
1994
Earlier work this paper cites.
——, “Approximation and estimation bounds for artificial neural networks,” Mach. Learn. , vol. 14, no. 1, pp. 115–133, 1994. [Online]. Available: http://dx.doi.org/10.1007/BF00993164
1994
Earlier work this paper cites.
C. K. Chui, X. Li, and H. N. Mhaskar, “Neural networks for localized approximation,” Math. Comp. , vol. 63, no. 208, pp. 607–623, 1994. [Online]. Available: http://dx.doi.org/10.2307/2153285
1994
Earlier work this paper cites.
S. Ellacott, “Aspects of the numerical analysis of neural networks,” Acta Numer. , vol. 3, pp. 145–202, 1994
1994
Earlier work this paper cites.
Y. LeCun, L. D. Jackel, L. Bottou, A. Brunot, C. Cortes, J. S. Denker, H. Drucker, I. Guyon, U. A. Müller, E. Säckinger, P. Simard, and V. Vapnik, “Comparison of learning algorithms for handwritten digit recognition,” International Conference on Artificial Neural Networks , pp. 53–60, 1995
1995
Earlier work this paper cites.
H. N. Mhaskar and C. A. Micchelli, “Degree of approximation by neural and translation networks with a single hidden layer,” Adv. Appl. Math. , vol. 16, no. 2, pp. 151–183, 1995
1995
Earlier work this paper cites.
——, “Unconditional bases and bit-level compression,” Appl. Comput. Harm. Anal. , vol. 3, pp. 388–392, 1996
1996
Earlier work this paper cites.
R. DeVore, K. Oskolkov, and P. Petrushev, “Approximation by feed-forward neural networks,” Ann. Numer. Math. , vol. 4, pp. 261–287, 1996
1996
Cited alongside, same era.
H. N. Mhaskar, “Neural networks for optimal approximation of smooth and analytic functions,” Neural Comput. , vol. 8, no. 1, pp. 164–177, 1996
1996
Cited alongside, same era.
M. Unser, “Ten good reasons for using spline wavelets,” Wavelet Applications in Signal and Image Processing V , vol. 3169, pp. 422–431, 1997
1997
Cited alongside, same era.
E. J. Candès, “Ridgelets: Theory and applications,” Ph.D. dissertation, Stanford University, 1998
1998
Cited alongside, same era.
R. A. DeVore, “Nonlinear approximation,” Acta Numerica , vol. 7, pp. 51–150, 1998
1998
Cited alongside, same era.
P. Grohs, S. Keiper, G. Kutyniok, and M. Schäfer, “ α \alpha -molecules,” Appl. Comput. Harmon. Anal. , vol. 41, no. 1, pp. 297–336, 2016. [Online]. Available: http://dx.doi.org/10.1016/j.acha.2015.10.009
2015
Later among the works it cites.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis, “Mastering the game of Go with deep neural networks and tree search,” Nature , vol. 529, no. 7587, pp. 484–489, 2016. [Online]. Available: http://www.nature.com/nature/journal/v529/n7587/abs/nature16961.html#supplementary-information
2016
Later among the works it cites.
I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning . MIT Press, 2016, http://www.deeplearningbook.org
2016
Later among the works it cites.
R. Eldan and O. Shamir, “The power of depth for feedforward neural networks,” in Proceedings of the 29th Conference on Learning Theory , 2016, pp. 907–940
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. L. Donoho, M. Vetterli, R. A. DeVore, and I. Daubechies, “Data compression and harmonic analysis,” IEEE Transactions on Information Theory , vol. 44, no. 6, pp. 2435–2476, 1998
1998
Cited alongside, same era.
M. Anthony and P. L. Bartlett, Neural Network Learning: Theoretical Foundations . Cambridge University Press, 1999
1999
Cited alongside, same era.
T. Nguyen-Thien and T. Tran-Cong, “Approximation of functions and their derivatives: A neural network implementation with applications,” Appl. Math. Model. , vol. 23, no. 9, pp. 687–704, 1999. [Online]. Available: //www.sciencedirect.com/science/article/pii/S0307904X99000062
1999
Cited alongside, same era.
A. Pinkus, “Approximation theory of the MLP model in neural networks,” Acta Numer. , vol. 8, pp. 143–195, 1999
1999
Cited alongside, same era.
K. Gröchenig and S. Samarah, “Nonlinear approximation with local Fourier bases,” Constructive Approximation , vol. 16, no. 3, pp. 317–331, Jul. 2000
2000
Cited alongside, same era.
J. Munkres, Topology , ser. Featured Titles for Topology. Prentice Hall, Incorporated, 2000
2000
Cited alongside, same era.
D. L. Donoho, “Sparse components of images and optimal atomic decompositions,” Constr. Approx. , vol. 17, no. 3, pp. 353–382, 2001. [Online]. Available: http://dx.doi.org/10.1007/s003650010032
2001
Cited alongside, same era.
2016
Later among the works it cites.
H. N. Mhaskar and T. Poggio, “Deep vs. shallow networks: An approximation theory perspective,” Analysis and Applications , vol. 14, no. 6, pp. 829–848, 2016. [Online]. Available: http://www.worldscientific.com/doi/abs/10.1142/S0219530516400042
2016
Later among the works it cites.
N. Cohen, O. Sharir, and A. Shashua, “On the expressive power of deep learning: A tensor analysis,” in Proceedings of the 29th Conference on Learning Theory , vol. 49, 2016, pp. 698–728
2016
Later among the works it cites.
N. Cohen and A. Shashua, “Convolutional rectifier networks as generalized tensor decompositions,” in Proceedings of the 33rd International Conference on Machine Learning , vol. 48, 2016, pp. 955–963
2016
Later among the works it cites.
P. Grohs, S. Keiper, G. Kutyniok, and M. Schäfer, “Cartoon approximation with α \alpha -curvelets,” J. Fourier Anal. Appl. , vol. 22, no. 6, pp. 1235–1293, 2016. [Online]. Available: http://dx.doi.org/10.1007/s00041-015-9446-6
2016
Later among the works it cites.
D. Yarotsky, “Error bounds for approximations with deep ReLU networks,” Neural Networks , vol. 94, pp. 103–114, 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
P. Petersen and F. Voigtlaender, “Optimal approximation of piecewise smooth functions using deep ReLU neural networks,” Neural Networks , vol. 108, pp. 296–330, Sep. 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
U. Shaham, A. Cloninger, and R. R. Coifman, “Provable approximation properties for deep neural networks,” Appl. Comput. Harmon. Anal. , vol. 44, no. 3, pp. 537–557, May 2018. [Online]. Available: http://dblp.uni-trier.de/db/journals/corr/corr1509.html#ShahamCC15
2018
Later among the works it cites.
M. Ehler and F. Filbir, “Metric entropy, n-widths, and sampling of functions on manifolds,” Journal of Approximation Theory , vol. 225, pp. 41 – 57, 2018
2018
Later among the works it cites.
H. Bölcskei, P. Grohs, G. Kutyniok, and P. Petersen, “Optimal approximation with sparsely connected deep neural networks,” SIAM Journal on Mathematics of Data Science , vol. 1, no. 1, pp. 8–45, 2019
2019
Closest in time.
B. Hanin and D. Rolnick, “Deep ReLU networks have surprisingly few activation patterns,” in Advances in Neural Information Processing Systems 32 . Curran Associates, Inc., 2019, pp. 361–370. [Online]. Available: http://papers.nips.cc/paper/8328-deep-relu-networks-have-surprisingly-few-activation-patterns.pdf
2019
Closest in time.
C. Schwab and J. Zech, “Deep learning in high dimension: Neural network expression rates for generalized polynomial chaos expansions in UQ,” Analysis and Applications , vol. 17, no. 1, pp. 19–55, 2019
2019
Closest in time.
M. Wainwright, High-dimensional statistics: A non-asymptotic viewpoint . Cambridge University Press, 2019
2019
Closest in time.
2019
Closest in time.
2020
Closest in time.
J. A. A. Opschoor, P. C. Petersen, and C. Schwab, “Deep ReLU networks and high-order finite element methods,” Analysis and Applications , vol. 18, no. 5, pp. 715–770, 2020. [Online]. Available: https://doi.org/10.1142/S0219530519410136
2020
Closest in time.
I. Gühring, G. Kutyniok, and P. Petersen, “Error bounds for approximations with deep ReLU neural networks in W s , p W^{s,p} norms,” Analysis and Applications , vol. 18, no. 5, pp. 803–859, 2020. [Online]. Available: https://doi.org/10.1142/S0219530519410021
2020
Closest in time.
J. Berner, P. Grohs, and A. Jentzen, “Analysis of the generalization error: Empirical risk minimization over deep artificial neural networks overcomes the curse of dimensionality in the numerical approximation of Black–Scholes partial differential equations,” SIAM Journal on Mathematics of Data Science , vol. 2, no. 3, pp. 631–657, 2020
2020
Closest in time.
R. DeVore, B. Hanin, and G. Petrova, “Neural network approximation,” arXiv:2012.14501 , 2020
2020
Closest in time.
H. Mhaskar, “A direct approach for function approximation on data defined manifolds,” Neural Networks , vol. 132, pp. 253 – 268, 2020
2020
Closest in time.
2020
Closest in time.
——, “Affine symmetries and neural network identifiability,” Advances in Mathematics , vol. 376, no. 107485, pp. 1–72, 2021. [Online]. Available: http://www.sciencedirect.com/science/article/pii/S0001870820305132
2021
Closest in time.