Fetching the paper…
Reading the bibliography…
Despite their immense promise in performing a variety of learning tasks, a theoretical understanding of the limitations of Deep Neural Networks (DNNs) has so far eluded practitioners.
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Brew, Huggingface’s transformers: State-of-the-art natural language processing , CoRR abs/1910.03771 · 1910
Earlier work this paper cites.
G. Cybenko, Approximation by superpositions of a sigmoidal function, Mathematics of control, signals and systems 2 (4) (1989) 303–314
1989
Earlier work this paper cites.
K. Hornik, M. Stinchcombe, H. White, Multilayer feedforward networks are universal approximators, Neural networks 2 (5) (1989) 359–366
1989
Earlier work this paper cites.
B. E. Boser, I. M. Guyon, V. N. Vapnik, A training algorithm for optimal margin classifiers, in: Proceedings of the fifth annual workshop on Computational learning theory, 1992, pp. 144–152
1992
Earlier work this paper cites.
B. Schölkopf, R. Herbrich, A. J. Smola, A generalized representer theorem, in: International conference on computational learning theory, Springer, 2001, pp. 416–426
2001
Earlier work this paper cites.
2003
Earlier work this paper cites.
A. Go, R. Bhayani, L. Huang, Twitter sentiment classification using distant supervision , in: Stanford CS224N Project Report, 2009. URL https://api.semanticscholar.org/CorpusID:18635269
2009
Earlier work this paper cites.
S. Fort, G. K. Dziugaite, M. Paul, S. Kharaghani, D. M. Roy, S. Ganguli, Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the neural tangent kernel (2020) · 2010
Earlier work this paper cites.
T. A. Driscoll, N. Hale, L. N. Trefethen, Chebfun Guide , Pafnuty Publications, 2014. URL http://www.chebfun.org/docs/guide/
2014
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778
2016
Cited alongside, same era.
D. Yarotsky, Optimal approximation of continuous functions by very deep ReLU networks, in: Conference on learning theory, PMLR, 2018, pp. 639–649
2018
Cited alongside, same era.
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, Improving language understanding by generative pre-training (2018). URL https://api.semanticscholar.org/CorpusID:49313245
2018
Cited alongside, same era.
J. D. M.-W. C. Kenton, L. K. Toutanova, Bert: Pre-training of deep bidirectional transformers for language understanding, in: Proceedings of naacL-HLT, Vol. 1, 2019, p. 2
2019
Cited alongside, same era.
L. Lu, P. Jin, G. Pang, Z. Zhang, G. E. Karniadakis, Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators, Nature machine intelligence 3 (3) (2021) 218–229
2021
Later among the works it cites.
N. Kovachki, S. Lanthaler, S. Mishra, On universal approximation and error bounds for Fourier neural operators, The Journal of Machine Learning Research 22 (1) (2021) 13237–13312
2021
Later among the works it cites.
P. L. Bartlett, A. Montanari, A. Rakhlin, Deep learning: a statistical viewpoint, Acta numerica 30 (2021) 87–201
2021
Later among the works it cites.
M. Belkin, Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation, Acta Numerica 30 (2021) 203–248
2021
Later among the works it cites.
S. Wang, H. Wang, P. Perdikaris, Learning the solution operator of parametric partial differential equations with physics-informed DeepONets, Science advances 7 (40) (2021) eabi8605
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
Z. Fan, Z. Wang, Spectra of the conjugate kernel and neural tangent kernel for linear-width neural networks, Advances in neural information processing systems 33 (2020) 7710–7721
2020
Cited alongside, same era.
G. Meanti, L. Carratino, L. Rosasco, A. Rudi, Kernel methods through the roof: handling billions of points efficiently, Advances in Neural Information Processing Systems 33 (2020) 14410–14422
2020
Cited alongside, same era.
C. Liu, L. Zhu, M. Belkin, On the linearity of large non-linear models: when and why the tangent kernel is constant, Advances in Neural Information Processing Systems 33 (2020) 15954–15964
2020
Cited alongside, same era.
G. E. Karniadakis, I. G. Kevrekidis, L. Lu, P. Perdikaris, S. Wang, L. Yang, Physics-informed machine learning, Nature Reviews Physics 3 (6) (2021) 422–440
2021
Cited alongside, same era.
A. Krizhevsky, I. Sutskever, G. E. Hinton, Imagenet classification with deep convolutional neural networks, Advances in neural information processing systems 25
Cited in the paper.
Cited in the paper.
Y. Kim, Convolutional neural networks for sentence classification, arXiv preprint arXiv:1408.5882
Cited in the paper.
2021
Later among the works it cites.
I. Daubechies, R. DeVore, S. Foucart, B. Hanin, G. Petrova, Nonlinear approximation and (deep) ReLU networks, Constructive Approximation 55 (1) (2022) 127–172
2022
Later among the works it cites.
R. Novak, J. Sohl-Dickstein, S. S. Schoenholz, Fast finite width neural tangent kernel, in: International Conference on Machine Learning, PMLR, 2022, pp. 17018–17044
2022
Later among the works it cites.
S. Wang, H. Wang, P. Perdikaris, Improved architectures and training algorithms for deep operator networks, Journal of Scientific Computing 92 (2) (2022) 35
2022
Later among the works it cites.
A. A. Howard, M. Perego, G. E. Karniadakis, P. Stinis, Multifidelity deep operator networks for data-driven and physics-informed problems, Journal of Computational Physics 493 (2023) 112462
2023
Closest in time.