Fetching the paper…
Reading the bibliography…
In this paper we analyze the $L_2$ error of neural network regression estimates with one hidden layer.
Kohler, M., and Langer, S. (2022). Discussion of “Nonparametric regression using deep neural networks with ReLU activation function”. Annals of Statistics
1910
Earlier work this paper cites.
Braun, A., Kohler, M., and Walk, H. (2019). On the rate of convergence of a neural network regression estimate learned by gradient descent. Preprint, arXiv: 1912.03921
1912
Earlier work this paper cites.
Andoni, A., Panigraphy, R., Valiant, G., and Zhang, L.(2014). Learning polynomials with neural networks. In International Conference on Machine Learning
1916
Earlier work this paper cites.
Yosida, K. (1968). Functional Analysis
1968
Earlier work this paper cites.
Pisier, G. "Remarques sur un resultat non publie de B. Maurey," presented at the Seminaire d’analyse fonctionelle 1980-1981, Ecole Polytechnique, Centre de Mathematiques, Palaiseau
1981
Earlier work this paper cites.
Stone, C. J. (1982). Optimal global rates of convergence for nonparametric regression. Annals of Statistics
1982
Earlier work this paper cites.
Barron, A. R. (1993). Universal approximation bounds for superpositions of a sigmoidal function. IEEE Transactions on Information Theory
1993
Earlier work this paper cites.
Barron, A. R. (1994). Approximation and estimation bounds for artificial neural networks. Machine Learning
1994
Earlier work this paper cites.
McCaffrey, D. F., and Gallant, A. R. (1994). Convergence rates for single hidden layer feedforward networks. Neural Networks
1994
Earlier work this paper cites.
Igelnik, B. and Pao, Y.-H. (1995). Stochastic choice of basis functions in adaptive function approximation and the functional-link net. IEEE Transactions on Neural Networks
1995
Earlier work this paper cites.
Lee, W.S. (1996). Agnostic Learning and Single Hidden Layer Neural Networks. PhD Thesis, Australian National University
1996
Earlier work this paper cites.
Györfi, L., Kohler, M., Krzyżak, A., and Walk, H. (2002). A Distribution–Free Theory of Nonparametric Regression
2002
Earlier work this paper cites.
Huang, G.-B., Chen, L., and Siew, C.-K. (2006). Universal approximation using incremental contractive feedforward networks with random hidden nodes. IEEE Transactions on Neural Networks
2006
Earlier work this paper cites.
Koltchinskii, V. (2006). Local Rademacher complexities and oracle inequalities in risk minimization. Annals of Statistics
2006
Earlier work this paper cites.
Epstein, Ch. L. (2008). Introduction to the Mathematics of Medical Imaging
2008
Earlier work this paper cites.
Kůrková, V., and Sanguinetti, M. (2008). Geometric upper bounds on rates of variable-basis approximation. IEEE Transactions on Information Theory
2008
Earlier work this paper cites.
Bagirov, A. M., Clausen, C., and Kohler, M. (2009). Estimation of a regression function by maxima of minima of linear functions. IEEE Transactions on Information Theory
2009
Earlier work this paper cites.
Rahimi, A., and Recht, B. (2009). Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning. In Advances in Neural Information Process Systems
2009
Earlier work this paper cites.
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012). ImageNet classification with deep convolutional neural networks. In Advances In Neural Information Processing Systems
2012
Cited alongside, same era.
Widrow, B., Greenblatt, A., Kim, Y., and Park, D. (2013). The No-Prop algorithm: A new learning algorithm for multilayer neural networks. Neural Networks
2013
Cited alongside, same era.
Kim, Y. (2014). Convolutional Neural Networks for Sentence Classification. In Empirical Methods in Natural Language Processing
2014
Cited alongside, same era.
Choromanska, A., Henaff, M., Mathieu, M., Arous, G. B., and LeCun, Y. (2015) The loss surface of multilayer networks. In International Conference on Artificial Intelligence and Statistics
2015
Cited alongside, same era.
Goodfellow, I., Bengio, Y., and Courville, A. (2016). Deep Learning
2016
Bauer, B., and Kohler, M. (2019). On deep learning as a remedy for the curse of dimensionality in nonparametric regression. Annals of Statistics
2019
Later among the works it cites.
Du, S., Lee, J., Li, H., Wang, L., und Zhai, X. (2019). Gradient descent finds global minima of deep neural networks. In International Conference on Machine Learning
2019
Later among the works it cites.
Dudek, G. (2019). Generating random weights and biases in feedfoward neural networks with random hidden nodes. Information Sciences
2019
Later among the works it cites.
Ghorbani, B., Mei, S., Misiakiewicz, T. and Montanari, A. (2019). Limitations of lazy training of two-layer neural networks. In Advances in Neural Information Processing Systems
2019
Later among the works it cites.
Kawaguchi, K, and Huang, J. (2019). Gradient descent finds global minima for generalizable deep neural networks of practical sizes. In 2019 57th annual allerton conference on communication, control, and computing (Allerton)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Kawaguchi, K. (2016). Deep learning without poor local minima. Advances in Neural Information Processing Systems
2016
Cited alongside, same era.
Wu, Y., Schuster, M., Chen, Z., Le, Q., Norouzi, M., Macherey, W., Krikum, M., et al. (2016). Google’s neural machine translation system: Bridging the gap between human and machine translation. arXiv: 1609.08144
2016
Cited alongside, same era.
Kohler, M., and Krzyżak, A. (2017). Nonparametric regression based on hierarchical interaction models. IEEE Transaction on Information Theory
2017
Cited alongside, same era.
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Huber, T., et al. (2017). Mastering the game of go without human knowledge. Nature
2017
Cited alongside, same era.
Arora, S., Cohen, N., Golowich, N., and Hu, W. (2018). A convergence analysis of gradient descent for deep linear neural networks. In International Conference on Learning Representations
2018
Cited alongside, same era.
Brutzkus, A., Globerson, A., Malach, E., and Shalew-Shwartz S. (2018). SGD learn overparametrized networks that provably generalize on linearly seperable data. In International Conference on Learning Representation
2018
Cited alongside, same era.
Du, S., and Lee, J. (2018). On the power of over-parametrization in neural networks with quadratic activation. In International Conference on Machine Learning
2018
Cited alongside, same era.
2019
Later among the works it cites.
Jacot, A., Gabriel, F., und Hongler, C. (2020). Neural Tangent Kernel: Convergence and Generalization in Neural Networks. In Advances in Neural Information Processing Systems, 31
2020
Later among the works it cites.
Schmidt-Hieber, J. (2020). Nonparametric regression using deep neural networks with ReLU activation function (with discussion). Annals of Statistics
2020
Later among the works it cites.
Sitzmann, V., Martel, J., Bergman,A., Lindell, D., and Wetzstein, G. (2020). Implicit neural representations with periodic activation functions. In Advances in Neural Information Processing Systems
2020
Later among the works it cites.
Woodworth, B., Gunasekar, S., Lee, J., Moroshko, E., Savarese, P., Golan, I., Soudry, D., und Srebro, N. (2020). Kernel and rich regimes in overparametrized models. In Conference on Learning Theory
2020
Later among the works it cites.
2020
Later among the works it cites.
Gonon, L. (2021). Random feature neural networks learn Black-Scholes type PDEs without curse of dimensionality. Preprint, arXiv: 2106.08900
2021
Closest in time.
Javanmard, A., Mondelli, M., and Montanari, A. (2021). Analysis of two-layer neural network via displacement convexity. Annals of Statistics
2021
Closest in time.
Kohler, M., and Krzyżak, A. (2021). Over-parametrized deep neural networks minimizing the empirical risk do not generalize well. Bernoulli
2021
Closest in time.
Sonoda, S., Ishikawa, I., and Ikeda, M. (2021). Ridge regression with overparametrized two-layer networks convergence to ridgelet spectrum. International Conference on Artificial Intelligence and Statistics
2021
Closest in time.
Suzuki, T., and Nitanda, A. (2021). Deep learning is adaptive to intrinsic diemsnionality of model smoothness in anisotropic Besov space. Advances in Neural Information Processing Systems
2021
Closest in time.
Kohler, M., Krzyżak, A., and Langer, S. (2022). Estimation of a function of low local dimensionality by deep neural networks. IEEE Transactions on Information Theory
2022
Closest in time.
Kohler, M., and Langer, S. (2022). On the rate of convergence of fully connected very deep neural network regression estimates using ReLU activation functions. Annals of Statistics
2022
Closest in time.
Gonon, L., Grigoryeva, L., and Ortega, J.-P. (2023). Approximation bounds for random neural networks and reservoir systems. The Annals of Applied Probability
2023
Closest in time.