Fetching the paper…
Reading the bibliography…
Modern machine learning systems such as deep neural networks are often highly over-parameterized so that they can fit the noisy training data exactly, yet they can still achieve small test errors in practice.
Two models of double descent for weak features
Belkin, M · 1903
Earlier work this paper cites.
Surprises in high-dimensional ridgeless least squares interpolation
Hastie, T · 1903
Earlier work this paper cites.
Gradient descent maximizes the margin of homogeneous neural networks
Lyu, K · 1906
Earlier work this paper cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
Mei, S · 1908
Earlier work this paper cites.
Finite-sample analysis of interpolating linear classifiers in the overparameterized regime
Chatterji, N. S · 2004
Earlier work this paper cites.
Classification vs regression in overparameterized regimes: Does the loss function matter?
Muthukumar, V · 2005
Earlier work this paper cites.
Montanari, A · 2007
Earlier work this paper cites.
Higher criticism thresholding: Optimal feature selection when useful features are rare and weak
Donoho, D · 2008
Earlier work this paper cites.
On the proliferation of support vectors in high dimensions
Hsu, D · 2009
Cited alongside, same era.
Impossibility of successful classification when useful features are rare and weak
Jin, J · 2009
Cited alongside, same era.
Benign overfitting in ridge regression
Tsigler, A · 2009
Cited alongside, same era.
Introduction to the non-asymptotic analysis of random matrices
Vershynin, R · 2010
Cited alongside, same era.
Benign overfitting in binary classification of gaussian mixtures
Wang, K · 2011
Cited alongside, same era.
Characterizing implicit bias in terms of optimization geometry
Gunasekar, S · 2018
Later among the works it cites.
The implicit bias of gradient descent on separable data
Soudry, D · 2018
Later among the works it cites.
Implicit regularization in deep matrix factorization
Arora, S · 2019
Later among the works it cites.
The implicit bias of gradient descent on nonseparable data
Ji, Z · 2019
Later among the works it cites.
Stochastic gradient descent on separable data: Exact convergence with a fixed learning rate
Nacson, M. S · 2019
Later among the works it cites.
Benign overfitting in linear regression
Bartlett, P. L · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A chernoff-type lower bound for the gaussian q-function
Côté, F. D · 2012
Cited alongside, same era.
Implicit regularization in matrix factorization
Gunasekar, S · 2017
Cited alongside, same era.
To understand deep learning we need to understand kernel learning
Belkin, M · 2018
Cited alongside, same era.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Belkin, M
Cited in the paper.
Harmless interpolation of noisy data in regression
Muthukumar, V
Cited in the paper.
A random matrix analysis of random fourier features: beyond the gaussian kernel, a precise phase transition, and the corresponding double descent
Liao, Z · 2020
Later among the works it cites.
On the optimal weighted ℓ 2 \ell_{2} regularization in overparameterized linear regression
Wu, D · 2020
Later among the works it cites.