Fetching the paper…
Reading the bibliography…
Interpolators -- estimators that achieve zero training error -- have attracted growing attention in machine learning, mainly because state-of-the art neural networks appear to be models of this type.
Linearized two-layers neural networks in high dimension
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 1904
Earlier work this paper cites.
Distribution of eigenvalues for some sets of random matrices
Vladimir Alexandrovich Marchenko and Leonid Andreevich Pastur · 1967
Earlier work this paper cites.
Smoothing noisy data with spline functions
Peter Craven and Grace Wahba · 1978
Earlier work this paper cites.
Smoothing noisy data with spline functions: Estimating the correct degree of smoothing by the method of generalized cross-validation
P. Craven and G. Wahba · 1979
Earlier work this paper cites.
Asymptotic optimality of C L C_{L} and generalized cross-validation in ridge regression with application to spline smoothing
Ker-Chau Li · 1986
Earlier work this paper cites.
Asymptotic optimality for C p C_{p} , C L C_{L} , cross-validation and generalized cross-validation: discrete index set
Ker-Chau Li · 1987
Earlier work this paper cites.
Limit of the smallest eigenvalue of a large dimensional sample covariance matrix
Zhidong Bai and Y. Q. Yin · 1993
Earlier work this paper cites.
Functional approximation by feed-forward networks: a least-squares approach to generalization
Andrew R Webb · 1994
Earlier work this paper cites.
Training with noise is equivalent to tikhonov regularization
Chris M Bishop · 1995
Earlier work this paper cites.
Strong convergence of the empirical distribution of eigenvalues of large dimensional random matrices
Jack Silverstein · 1995
Earlier work this paper cites.
Finite Difference and Spectral Methods for Ordinary and Partial Differential Equations
Lloyd N. Trefethen · 1996
Earlier work this paper cites.
Asymptotic distribution of the spectra of a class of generalized Kac-Murdock-Szego matrices
William F. Trench · 1999
Earlier work this paper cites.
The entire regularization path for the support vector machine
Trevor Hastie, Saharon Rosset, Robert Tibshirani, and Ji Zhu · 2004
Earlier work this paper cites.
Boosting as a regularized path to a maximum margin classifier
Saharon Rosset, Ji Zhu, and Trevor Hastie · 2004
Earlier work this paper cites.
Random matrix theory and wireless communications
Antonia M. Tulino and Sergio Verdu · 2004
Earlier work this paper cites.
Multiparametric Statistics
Vadim I. Serdobolskii · 2007
Earlier work this paper cites.
An Introduction to Random Matrices
Greg W. Anderson, Alice Guionnet, and Ofer Zeitouni · 2009
Earlier work this paper cites.
A survey of cross-validation procedures for model selection
Sylvain Arlot and Alain Celisse · 2010
Earlier work this paper cites.
Spectral Analysis of Large Dimensional Random Matrices
Zhidong Bai and Jack Silverstein · 2010
Earlier work this paper cites.
The spectrum of kernel random matrices
Noureddine El Karoui · 2010
Earlier work this paper cites.
Eigenvectors of some large sample covariance matrix ensembles
Olivier Ledoit and Sandrine Peche · 2011
Cited alongside, same era.
Spectral convergence for a general class of random matrices
Francisco Rubio and Xavier Mestre · 2011
Cited alongside, same era.
Topics in Random Matrix Theory , volume 132
Terence Tao · 2012
Cited alongside, same era.
The spectrum of random inner-product kernel matrices
Xiuyuan Cheng and Amit Singer · 2013
Cited alongside, same era.
Efficient computation of limit spectra of sample covariance matrices
Edgar Dobriban · 2015
Cited alongside, same era.
The spectral norm of random inner-product kernel matrices
Zhou Fan and Andrea Montanari · 2015
Cited alongside, same era.
Mean field analysis of neural networks
Justin Sirignano and Konstantinos Spiliopoulos · 2018
Later among the works it cites.
A jamming transition from under-to over-parametrization affects loss landscape and generalization
Stefano Spigler, Mario Geiger, Stephane d’Ascoli, Levent Sagun, Giulio Biroli, and Matthieu Wyart · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep ReLU networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Later among the works it cites.
A continuous-time view of early stopping for least squares
Alnur Ali, J. Zico Kolter, and Ryan J. Tibshirani · 2019
Closest in time.
Two models of double descent for weak features
Mikhail Belkin, Daniel Hsu, and Ji Xu · 2019
Closest in time.
Scaling description of generalization with number of parameters in deep learning
Mario Geiger, Arthur Jacot, Stefano Spigler, Franck Gabriel, Levent Sagun, Stephane d’Ascoli, Giulio Biroli, Clement Hongler, and Matthieu Wyart · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A general framework for fast stagewise algortihms
Ryan J. Tibshirani · 2015
Cited alongside, same era.
Ridge regression and asymptotic minimax estimation over spheres of growing dimension
Lee H. Dicker · 2016
Cited alongside, same era.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
Madhu S. Advani and Andrew M. Saxe · 2017
Cited alongside, same era.
Anisotropic local laws for random matrices
Antti Knowles and Jun Yin · 2017
Cited alongside, same era.
Closest in time.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel S Schoenholz, Yasaman Bahri, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Closest in time.
The generalization error of random features regression: Precise asymptotics and double descent curve
Song Mei and Andrea Montanari · 2019
Closest in time.
Andrea Montanari, Feng Ruan, Youngtak Sohn, and Jun Yan · 2019
Closest in time.
Consistent risk estimation in high-dimensional linear regression
Ji Xu, Arian Maleki, and Kamiar Rahnama Rad · 2019
Closest in time.
Are all layers created equal?
Chiyuan Zhang, Samy Bengio, and Yoram Singer · 2019
Closest in time.
Ben Adlam and Jeffrey Pennington · 2020
Closest in time.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2020
Closest in time.
Generalisation error in learning with random features and the hidden manifold model
Federica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard, and Lenka Zdeborová · 2020
Closest in time.
The gaussian equivalence of generative models for learning with two-layer neural networks
Sebastian Goldt, Galen Reeves, Marc Mezard, Florent Krzakala, and Lenka Zdeborová · 2020
Closest in time.
Universality laws for high-dimensional learning with random features
Hong Hu and Yue M Lu · 2020
Closest in time.
The optimal ridge penalty for real-world high-dimensional data can be zero or negative due to the implicit ridge regularization
Dmitry Kobak, Jonathan Lomond, and Benoit Sanchez · 2020
Closest in time.
Just interpolate: Kernel “ridgeless” regression can generalize
Tengyuan Liang, Alexander Rakhlin, et al · 2020
Closest in time.
Asymptotics of ridge (less) regression under general source condition
Dominic Richards, Jaouad Mourtada, and Lorenzo Rosasco · 2020
Closest in time.
On the optimal weighted ℓ 2 \ell_{2} regularization in overparameterized linear regression
Denny Wu and Ji Xu · 2020
Closest in time.