Fetching the paper…
Reading the bibliography…
The classical bias-variance trade-off predicts that bias decreases and variance increase with model complexity, leading to a U-shaped risk curve.
A limit theorem for the norm of random matrices
Geman, S · 1980
Earlier work this paper cites.
Neural networks and the bias/variance dilemma
Geman, S., Bienenstock, E., and Doursat, R · 1992
Earlier work this paper cites.
A simple weight decay can improve generalization
Krogh, A. and Hertz, J. A · 1992
Earlier work this paper cites.
Machine learning bias, statistical bias, and statistical variance of decision tree algorithms
Dietterich, T. G. and Kong, E. B · 1995
Earlier work this paper cites.
The mnist database of handwritten digits
LeCun, Y · 1998
Earlier work this paper cites.
The Elements of Statistical Learning
Hastie, T., Tibshirani, R., and Friedman, J · 2001
Earlier work this paper cites.
Inversion error, condition number, and approximate inverses of uncertain matrices
Ghaoui, L. E · 2002
Earlier work this paper cites.
Pattern Recognition and Machine Learning
Bishop, C. M · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images, 2009
Krizhevsky, A. et al · 2009
Earlier work this paper cites.
Spectral Analysis of Large Dimensional Random Matrices
Bai, Z. and Silverstein, J · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
A generalized bias-variance decomposition for bregman divergences, 2013
Pfau, D · 2013
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
Advani, M. S. and Saxe, A. M · 2017
Cited alongside, same era.
Rosset, S. and Tibshirani, R. J · 2017
Cited alongside, same era.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Xiao, H., Rasul, K., and Vollgraf, R · 2017
Cited alongside, same era.
Surprises in High-Dimensional Ridgeless Least Squares Interpolation
Hastie, T., Montanari, A., Rosset, S., and Tibshirani, R. J · 2019
Later among the works it cites.
Benchmarking neural network robustness to common corruptions and perturbations
Hendrycks, D. and Dietterich, T · 2019
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
Mei, S. and Montanari, A · 2019
Later among the works it cites.
More data can hurt for linear regression: Sample-wise double descent
Nakkiran, P · 2019
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aggregated residual transformations for deep neural networks
Xie, S., Girshick, R., Dollár, P., Tu, Z., and He, K · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2017
Cited alongside, same era.
To understand deep learning we need to understand kernel learning
Belkin, M., Ma, S., and Mandal, S · 2018
Cited alongside, same era.
An Introduction to Wishart Matrix Moments
Bishop, A. N., Del Moral, P., and Niclas, A · 2018
Cited alongside, same era.
Why do deep convolutional networks generalize so poorly to small image transformations?
Azulay, A. and Weiss, Y · 2019
Cited alongside, same era.
A model of double descent for high-dimensional binary linear classification
Deng, Z., Kammoun, A., and Thrampoulidis, C · 2019
Cited alongside, same era.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Belkin, M., Hsu, D., Ma, S., and Mandal, S
Cited in the paper.
Later among the works it cites.
A modern take on the bias-variance tradeoff in neural networks, 2019
Neal, B., Mittal, S., Baratin, A., Tantia, V., Scicluna, M., Lacoste-Julien, S., and Mitliagkas, I · 2019
Later among the works it cites.
A jamming transition from under-to over-parametrization affects generalization in deep learning
Spigler, S., Geiger, M., d’Ascoli, S., Sagun, L., Biroli, G., and Wyart, M · 2019
Later among the works it cites.
High-Dimensional Statistics: A Non-Asymptotic Viewpoint
Wainwright, M. J · 2019
Later among the works it cites.
Generalization of two-layer neural networks: An asymptotic viewpoint
Ba, J., Erdogdu, M., Suzuki, T., Wu, D., and Zhang, T · 2020
Closest in time.
Benign overfitting in linear regression
Bartlett, P. L., Long, P. M., Lugosi, G., and Tsigler, A · 2020
Closest in time.
Finite-sample analysis of interpolating linear classifiers in the overparameterized regime, 2020
Chatterji, N. S. and Long, P. M · 2020
Closest in time.