Fetching the paper…
Reading the bibliography…
Recent advances in machine learning have been achieved by using overparametrized models trained until near interpolation of the training data.
Gabor Szeg, Orthogonal polynomials , vol. 23, American Mathematical Soc., 1939
1939
Earlier work this paper cites.
Peter Craven and Grace Wahba, Smoothing noisy data with spline functions: estimating the correct degree of smoothing by the method of generalized cross-validation , Numerische mathematik 31
1978
Earlier work this paper cites.
Gene H Golub, Michael Heath, and Grace Wahba, Generalized cross-validation as a method for choosing a good ridge parameter , Technometrics 21
1979
Earlier work this paper cites.
William Beckner, Sobolev inequalities, the poisson semigroup, and analysis on the sphere sn. , Proceedings of the National Academy of Sciences 89
1992
Earlier work this paper cites.
Andrew R Barron, Universal approximation bounds for superpositions of a sigmoidal function , IEEE Transactions on Information theory 39
1993
Earlier work this paper cites.
Radford M Neal, Bayesian learning for neural networks , Ph.D. thesis, Citeseer, 1995
1995
Earlier work this paper cites.
Hrushikesh N Mhaskar, Neural networks for optimal approximation of smooth and analytic functions , Neural computation 8
1996
Earlier work this paper cites.
Vitaly E Maiorov, On best approximation by ridge functions , Journal of Approximation Theory 99
1999
Earlier work this paper cites.
Allan Pinkus, Approximation theory of the mlp model in neural networks , Acta numerica 8
1999
Earlier work this paper cites.
Vladimir N Vapnik, An overview of statistical learning theory , IEEE transactions on neural networks 10
1999
Earlier work this paper cites.
Andrew Ng, Cs229 lecture notes , CS229 Lecture notes 1
2000
Earlier work this paper cites.
Guang-Bin Huang, Qin-Yu Zhu, and Chee-Kheong Siew, Extreme learning machine: theory and applications , Neurocomputing 70
2006
Earlier work this paper cites.
Léon Bottou and Olivier Bousquet, The tradeoffs of large scale learning , Advances in neural information processing systems 20
2007
Earlier work this paper cites.
Andrea Caponnetto and Ernesto De Vito, Optimal rates for the regularized least-squares algorithm , Foundations of Computational Mathematics 7
2007
Earlier work this paper cites.
Ali Rahimi and Benjamin Recht, Random features for large-scale kernel machines , Advances in neural information processing systems, 2008, pp. 1177–1184
2008
Earlier work this paper cites.
Noureddine El Karoui et al., The spectrum of kernel random matrices , The Annals of Statistics 38
2010
Earlier work this paper cites.
Theodore S Chihara, An introduction to orthogonal polynomials , Courier Corporation, 2011
2011
Earlier work this paper cites.
Joel A Tropp, User-friendly tail bounds for sums of random matrices , Foundations of computational mathematics 12
2012
Earlier work this paper cites.
Xiuyuan Cheng and Amit Singer, The spectrum of random inner-product kernel matrices , Random Matrices: Theory and Applications 2
2013
Earlier work this paper cites.
Feng Dai and Yuan Xu, Approximation theory and harmonic analysis on spheres and balls , vol. 23, Springer, 2013
2013
Earlier work this paper cites.
Francis Bach, Breaking the curse of dimensionality with convex neural networks , The Journal of Machine Learning Research 18
2017
Earlier work this paper cites.
László Erdős and Horng-Tzer Yau, A dynamical approach to random matrix theory , vol. 28, American Mathematical Soc., 2017
2017
Earlier work this paper cites.
Jeffrey Pennington and Pratik Worah, Nonlinear random matrix theory for deep learning , Advances in Neural Information Processing Systems, 2017, pp. 2637–2646
2017
Cited alongside, same era.
Alessandro Rudi and Lorenzo Rosasco, Generalization properties of learning with random features , Advances in neural information processing systems 30
2017
Cited alongside, same era.
2018
Cited alongside, same era.
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh, Gradient descent provably optimizes over-parameterized neural networks , International Conference on Learning Representations, 2018
2018
Cited alongside, same era.
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari, When do neural networks outperform kernel methods? , Advances in Neural Information Processing Systems 33
2020
Later among the works it cites.
Zhenyu Liao, Romain Couillet, and Michael W Mahoney, A random matrix analysis of random fourier features: beyond the gaussian kernel, a precise phase transition, and the corresponding double descent , Advances in Neural Information Processing Systems 33
2020
Later among the works it cites.
Tengyuan Liang, Alexander Rakhlin, et al., Just interpolate: Kernel “ridgeless” regression can generalize , Annals of Statistics 48
2020
Later among the works it cites.
Vidya Muthukumar, Kailas Vodrahalli, Vignesh Subramanian, and Anant Sahai, Harmless interpolation of noisy data in regression , IEEE Journal on Selected Areas in Information Theory 1
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
Yuanzhi Li and Yingyu Liang, Learning overparameterized neural networks via stochastic gradient descent on structured data , Advances in Neural Information Processing Systems, 2018, pp. 8157–8166
2018
Cited alongside, same era.
Cosme Louart, Zhenyu Liao, and Romain Couillet, A random matrix approach to neural networks , The Annals of Applied Probability 28
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Roman Vershynin, High-dimensional probability: An introduction with applications in data science , Cambridge University Press, 2018
2018
Cited alongside, same era.
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang, Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks , International Conference on Machine Learning, PMLR, 2019, pp. 322–332
2019
Cited alongside, same era.
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song, A convergence theory for deep learning via over-parameterization , International Conference on Machine Learning, PMLR, 2019, pp. 242–252
2019
Cited alongside, same era.
2020
Later among the works it cites.
Johannes Schmidt-Hieber, Nonparametric regression using deep neural networks with relu activation function , Annals of statistics 48
2020
Later among the works it cites.
Peter L Bartlett, Andrea Montanari, and Alexander Rakhlin, Deep learning: a statistical viewpoint , Acta numerica 30
2021
Later among the works it cites.
2021
Later among the works it cites.
Ben Adlam, Jake A Levinson, and Jeffrey Pennington, A random matrix perspective on mixtures of nonlinearities in high dimensions , International Conference on Artificial Intelligence and Statistics, PMLR, 2022, pp. 3434–3457
2022
Later among the works it cites.
Romain Couillet and Zhenyu Liao, Random matrix methods for machine learning , Cambridge University Press, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
Sebastian Goldt, Bruno Loureiro, Galen Reeves, Florent Krzakala, Marc Mézard, and Lenka Zdeborová, The gaussian equivalence of generative models for learning with shallow neural networks , Mathematical and Scientific Machine Learning, PMLR, 2022, pp. 426–471
2022
Later among the works it cites.
2022
Later among the works it cites.
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani, Surprises in high-dimensional ridgeless least squares interpolation , The Annals of Statistics 50
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Song Mei and Andrea Montanari, The generalization error of random features regression: Precise asymptotics and the double descent curve , Communications on Pure and Applied Mathematics 75
2022
Later among the works it cites.
Song Mei, Theodor Misiakiewicz, and Andrea Montanari, Generalization error of random feature and kernel methods: hypercontractivity and kernel matrix concentration , Applied and Computational Harmonic Analysis 59
2022
Later among the works it cites.
Andrea Montanari and Basil N Saeed, Universality of empirical risk minimization , Conference on Learning Theory, PMLR, 2022, pp. 4310–4312
2022
Later among the works it cites.
Andrea Montanari and Yiqiao Zhong, The interpolation phase transition in neural networks: Memorization and generalization under lazy training , The Annals of Statistics 50
2022
Later among the works it cites.
Lechao Xiao, Hong Hu, Theodor Misiakiewicz, Yue Lu, and Jeffrey Pennington, Precise learning curves and higher-order scalings for dot-product kernel regression , Advances in Neural Information Processing Systems 35
2022
Later among the works it cites.