Fetching the paper…
Reading the bibliography…
Recent empirical and theoretical studies have shown that many learning algorithms -- from linear regression to neural networks -- can have test performance that is non-monotonic in quantities such the sample size and model size.
Scaling description of generalization with number of parameters in deep learning
Geiger, M., Jacot, A., Spigler, S., Gabriel, F., Sagun, L., d’Ascoli, S., Biroli, G., Hongler, C., and Wyart, M · 1901
Earlier work this paper cites.
Improved sample complexities for deep networks and robust classification via an all-layer margin
Wei, C. and Ma, T · 1910
Earlier work this paper cites.
A problem of dimensionality: A simple example
Trunk, G. V · 1979
Earlier work this paper cites.
Matrix Analysis
Horn, R. A., Horn, R. A., and Johnson, C. R · 1990
Earlier work this paper cites.
Eigenvalues of covariance matrices: Application to neural-network learning
Le Cun, Y., Kanter, I., and Solla, S. A · 1991
Earlier work this paper cites.
Second order properties of error surfaces: Learning time and generalization
LeCun, Y., Kanter, I., and Solla, S. A · 1991
Earlier work this paper cites.
Small sample size generalization
Duin, R. P · 1995
Earlier work this paper cites.
Statistical mechanics of learning: Generalization
Opper, M · 1995
Earlier work this paper cites.
Classifiers in almost empty spaces
Duin, R. P · 2000
Earlier work this paper cites.
Learning to generalize
Opper, M · 2001
Earlier work this paper cites.
Bagging, boosting and the random subspace method for linear classifiers
Skurichina, M. and Duin, R. P · 2002
Earlier work this paper cites.
Random features for large-scale kernel machines
Rahimi, A. and Recht, B · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
The dipping phenomenon
Loog, M. and Duin, R. P · 2012
Earlier work this paper cites.
Ramanujan graphs and the solution of the kadison-singer problem
Marcus, A. W., Spielman, D. A., and Srivastava, N · 2014
Earlier work this paper cites.
High-dimensional dynamics of generalization error in neural networks
Advani, M. S. and Saxe, A. M · 2017
Cited alongside, same era.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Xiao, H., Rasul, K., and Vollgraf, R · 2017
Cited alongside, same era.
Reconciling modern machine learning and the bias-variance trade-off
Belkin, M., Hsu, D., Ma, S., and Mandal, S · 2018
Cited alongside, same era.
High-dimensional asymptotics of prediction: Ridge regression and classification
Dobriban, E., Wager, S., et al · 2018
Cited alongside, same era.
Wonder: Weighted one-shot distributed ridge regression in high dimensions
Dobriban, E. and Sheng, Y · 2019
Later among the works it cites.
Surprises in high-dimensional ridgeless least squares interpolation, 2019
Hastie, T., Montanari, A., Rosset, S., and Tibshirani, R. J · 2019
Later among the works it cites.
On the risk of minimum-norm interpolants and restricted lower isometry of kernels
Liang, T., Rakhlin, A., and Zhai, X · 2019
Later among the works it cites.
Minimizers of the empirical risk and risk monotonicity
Loog, M., Viering, T., and Mey, A · 2019
Later among the works it cites.
Asymptotic risk of least squares minimum norm estimator under the spike covariance model
Mahdaviyeh, Y. and Naulet, Z · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kobak, D., Lomond, J., and Sanchez, B · 2018
Cited alongside, same era.
An analytic theory of generalization dynamics and transfer learning in deep linear networks
Lampinen, A. K. and Ganguli, S · 2018
Cited alongside, same era.
Just interpolate: Kernel” ridgeless” regression can generalize
Liang, T. and Rakhlin, A · 2018
Cited alongside, same era.
A modern take on the bias-variance tradeoff in neural networks
Neal, B., Mittal, S., Baratin, A., Tantia, V., Scicluna, M., Lacoste-Julien, S., and Mitliagkas, I · 2018
Cited alongside, same era.
How to train your resnet
Page, D · 2018
Cited alongside, same era.
A jamming transition from under-to over-parametrization affects loss landscape and generalization
Spigler, S., Geiger, M., d’Ascoli, S., Sagun, L., Biroli, G., and Wyart, M · 2018
Cited alongside, same era.
Benign overfitting in linear regression
Bartlett, P. L., Long, P. M., Lugosi, G., and Tsigler, A · 2019
Cited alongside, same era.
Two models of double descent for weak features
Belkin, M., Hsu, D., and Xu, J · 2019
Cited alongside, same era.
The generalization error of random features regression: Precise asymptotics and double descent curve
Mei, S. and Montanari, A · 2019
Later among the works it cites.
Mitra, P. P · 2019
Later among the works it cites.
Harmless interpolation of noisy data in regression
Muthukumar, V., Vodrahalli, K., and Sahai, A · 2019
Later among the works it cites.
More data can hurt for linear regression: Sample-wise double descent
Nakkiran, P · 2019
Later among the works it cites.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
Wei, C., Lee, J. D., Liu, Q., and Ma, T · 2019
Later among the works it cites.
On the number of variables to use in principal component regression
Xu, J. and Hsu, D. J · 2019
Later among the works it cites.
Triple descent and the two kinds of overfitting: Where & why do they appear?
d’Ascoli, S., Sagun, L., and Biroli, G · 2020
Closest in time.
On the multiple descent of minimum-norm interpolants and restricted lower isometry of kernels
Liang, T., Rakhlin, A., and Zhai, X · 2020
Closest in time.
Deep double descent: Where bigger models and more data hurt
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I · 2020
Closest in time.