Fetching the paper…
Reading the bibliography…
Covering numbers of (deep) ReLU networks have been used to characterize approximation-theoretic performance, to upper-bound prediction error in nonparametric regression, and to quantify classification capacity.
V. Vapnik and A. Chervonenkis, “On the uniform convergence of relative frequencies of events to their probabilities,” Theory Probab. Appl. , vol. 16, no. 2, pp. 264–280, 1971
1971
Earlier work this paper cites.
B. M. Sh and M. Solomjak, Quantitative analysis in Sobolev imbedding theorems and applications to spectral theory , ser. American Mathematical Society translations. American Mathematical Society, 1980
1980
Earlier work this paper cites.
C. J. Stone, “Optimal global rates of convergence for nonparametric regression,” The Annals of Statistics , vol. 10, no. 4, pp. 1040 – 1053, Dec 1982
1982
Earlier work this paper cites.
K. Hornik, M. Stinchcombe, and H. White, “Multilayer feedforward networks are universal approximators,” Neural Networks , vol. 2, no. 5, pp. 359–366, Jan 1989
1989
Earlier work this paper cites.
G. Cybenko, “Approximation by superpositions of a sigmoidal function,” Mathematics of Control, Signals, and Systems , vol. 2, no. 4, pp. 303–314, Dec 1989
1989
Earlier work this paper cites.
K.-I. Funahashi, “On the approximate realization of continuous mappings by neural networks,” Neural Networks , vol. 2, no. 3, pp. 183–192, Jan 1989
1989
Earlier work this paper cites.
S. A. Janowsky, “Pruning versus clipping in neural networks,” Physical Review A , vol. 39, no. 12, pp. 6600–6603, June 1989
1989
Earlier work this paper cites.
A. Barron, “Universal approximation bounds for superpositions of a sigmoidal function,” IEEE Transactions on Information Theory , vol. 39, no. 3, pp. 930–945, May 1993
1993
Earlier work this paper cites.
D. L. Donoho, “Unconditional bases are optimal bases for data compression and for statistical estimation,” Applied and Computational Harmonic Analysis , vol. 1, no. 1, pp. 100–115, 1993
1993
Earlier work this paper cites.
——, “Unconditional bases and bit-level compression,” Applied and Computational Harmonic Analysis , vol. 3, no. 4, pp. 388–392, 1996
1996
Earlier work this paper cites.
M. Anthony and P. L. Bartlett, Neural Network Learning: Theoretical Foundations . Cambridge University Press, 1999, vol. 9
1999
Earlier work this paper cites.
Y. Yang and A. Barron, “Information-theoretic determination of minimax rates of convergence,” Annals of Statistics , vol. 27, no. 4, pp. 1564–1599, 1999
1999
Earlier work this paper cites.
L. Györfi, M. Kohler, A. Krzyżak, and H. Walk, A distribution-free theory of nonparametric regression , ser. Springer Series in Statistics. Springer, 2002, vol. 1
2002
Earlier work this paper cites.
2003
Earlier work this paper cites.
P. Grohs, “Optimally sparse data representations,” in Harmonic and Applied Analysis: From Groups to Signals . Springer, 2015, pp. 199–248
2015
Cited alongside, same era.
M. Telgarsky, “Benefits of depth in neural networks,” in 29th Annual Conference on Learning Theory , ser. Proceedings of Machine Learning Research, vol. 49, Columbia University, New York, USA, 23–26 June 2016, pp. 1517–1539
2016
Cited alongside, same era.
P. L. Bartlett, D. J. Foster, and M. J. Telgarsky, “Spectrally-normalized margin bounds for neural networks,” Advances in Neural Information Processing Systems , vol. 30, 2017
2017
Cited alongside, same era.
D. Yarotsky, “Error bounds for approximations with deep ReLU networks,” Neural Networks , vol. 94, pp. 103 – 114, 2017
2017
Cited alongside, same era.
P. Petersen and F. Voigtlaender, “Optimal approximation of piecewise smooth functions using deep ReLU neural networks,” Neural Networks , vol. 108, pp. 296–330, 2018
D. Blalock, J. J. Gonzalez Ortiz, J. Frankle, and J. Guttag, “What is the state of neural network pruning?” Proceedings of Machine Learning and Systems , vol. 2, pp. 129–146, 2020
2020
Later among the works it cites.
M. Kohler and S. Langer, “On the rate of convergence of fully connected deep neural network regression estimates,” The Annals of Statistics , vol. 49, no. 4, pp. 2231 – 2249, 2021
2021
Later among the works it cites.
D. Elbrächter, D. Perekrestenko, P. Grohs, and H. Bölcskei, “Deep neural network approximation theory,” IEEE Transactions on Information Theory , vol. 67, no. 5, pp. 2581–2623, May 2021
2021
Later among the works it cites.
I. Gühring and M. Raslan, “Approximation rates for neural networks with encodable weights in smoothness spaces,” Neural Networks , vol. 134, pp. 107–130, 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
M. J. Wainwright, High-Dimensional Statistics: A Non-Asymptotic Viewpoint , 2nd ed. Cambridge University Press, 2019, vol. 48
2019
Cited alongside, same era.
P. L. Bartlett, N. Harvey, C. Liaw, and A. Mehrabian, “Nearly-tight VC-dimension and pseudodimension bounds for piecewise linear neural networks.” Journal of Machine Learning Research , vol. 20, no. 63, pp. 1–17, 2019
2019
Cited alongside, same era.
H. Bölcskei, P. Grohs, G. Kutyniok, and P. Petersen, “Optimal approximation with sparsely connected deep neural networks,” SIAM Journal on Mathematics of Data Science , vol. 1, no. 1, pp. 8–45, 2019
2019
Cited alongside, same era.
Z. Shen, H. Yang, and S. Zhang, “Nonlinear approximation via compositions,” Neural Networks , vol. 119, pp. 74–84, 2019
2019
Cited alongside, same era.
Z. Shen, H. Yang, and S. Zhang, “Deep network approximation characterized by number of neurons,” Communications in Computational Physics , vol. 28, no. 5, pp. 1768–1811, 2020
2020
Cited alongside, same era.
D. Yarotsky and A. Zhevnerchuk, “The phase diagram of approximation rates for deep neural networks,” in Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 13 005–13 015
2020
Cited alongside, same era.
J. Schmidt-Hieber, “Nonparametric regression using deep neural networks with ReLU activation function,” The Annals of Statistics , vol. 48, no. 4, pp. 1875 – 1897, 2020
2020
Cited alongside, same era.
J. Gou, B. Yu, S. J. Maybank, and D. Tao, “Knowledge distillation: A survey,” International Journal of Computer Vision , vol. 129, no. 6, pp. 1789–1819, 2021
2021
Later among the works it cites.
A. Caragea, D. G. Lee, J. Maly, G. Pfander, and F. Voigtlaender, “Quantitative approximation results for complex-valued neural networks,” SIAM Journal on Mathematics of Data Science , vol. 4, no. 2, pp. 553–580, 2022
2022
Later among the works it cites.
M. Chen, H. Jiang, W. Liao, and T. Zhao, “Nonparametric regression on low-dimensional manifolds using deep ReLU networks: Function approximation and statistical recovery,” Information and Inference: A Journal of the IMA , vol. 11, no. 4, pp. 1203–1253, Mar. 2022
2022
Later among the works it cites.
D. Elbrächter, P. Grohs, A. Jentzen, and C. Schwab, “DNN expression rate analysis of high-dimensional PDEs: Application to option pricing,” Constructive Approximation , vol. 55, no. 1, pp. 3–71, 2022
2022
Later among the works it cites.
C. Huyen, Designing Machine Learning Systems . O’Reilly Media, Inc., 2022
2022
Later among the works it cites.
G. Vardi, G. Yehudai, and O. Shamir, “Width is less important than depth in ReLU neural networks,” in Proc. of Thirty Fifth Conference on Learning Theory , ser. Proceedings of Machine Learning Research, vol. 178. PMLR, July 2022, pp. 1249–1281
2022
Later among the works it cites.
A. Caragea, P. Petersen, and F. Voigtlaender, “Neural network approximation and estimation of classifiers with classification boundary in a Barron class,” The Annals of Applied Probability , vol. 33, no. 4, pp. 3039 – 3079, 2023
2023
Later among the works it cites.
J. Maly and R. Saab, “A simple approach for quantizing neural networks,” Applied and Computational Harmonic Analysis , vol. 66, pp. 138–150, 2023
2023
Later among the works it cites.