Fetching the paper…
Reading the bibliography…
Deep neural networks (DNNs) defy the classical bias-variance trade-off: adding parameters to a DNN that interpolates its training data will typically improve its generalization performance.
Neural networks and the bias/variance dilemma
S. Geman, E. Bienenstock, and R. Doursat · 1992
Earlier work this paper cites.
The MNIST database of handwritten digits, 1998
Y. LeCun and C. Cortes · 1998
Earlier work this paper cites.
The Elements of Statistical Learning
T. Hastie, R. Tibshirani, and J. Friedman · 2001
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky, G. Hinton, et al · 2009
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
G. E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2012
Earlier work this paper cites.
Representation learning: A review and new perspectives
Y. Bengio, A. Courville, and P. Vincent · 2013
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
B. Neyshabur, R. Tomioka, and N. Srebro · 2015
Earlier work this paper cites.
S. Zagoruyko and N. Komodakis · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Earlier work this paper cites.
A Closer Look at Memorization in Deep Networks
D. Arpit, S. Jastrzebski, M.S. Kanwal, T. Maharaj, A. Fischer, A. Courville, and Y. Bengio · 2017
Earlier work this paper cites.
Densely connected convolutional networks
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger · 2017
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
J. Hestness, S. Narang, N. Ardalani, G. Diamos, H. Jun, H. Kianinejad, M. Patwary, M. Ali, Y. Yang, and Y. Zhou · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Earlier work this paper cites.
Characterizing implicit bias in terms of optimization geometry
S. Gunasekar, J. Lee, D. Soudry, and N. Srebro · 2018
Earlier work this paper cites.
Implicit bias of gradient descent on linear convolutional networks
S. Gunasekar, J. D. Lee, D. Soudry, and N. Srebro · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
D. Soudry, E. Hoffer, and N. Srebro · 2018
Earlier work this paper cites.
A modern take on the bias-variance tradeoff in neural networks
B. Neal, S. Mittal, A. Baratin, V. Tantia, M. Scicluna, S. Lacoste-Julien, and I. Mitliagkas · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Cited alongside, same era.
Gradient descent learns one-hidden-layer CNN: Don’t be afraid of spurious local minima
S. Du, J. Lee, Y. Tian, A. Singh, and B. Poczos · 2018
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
S. Mei, A. Montanari, and P. Nguyen · 2018
Cited alongside, same era.
Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks
G.M. Rotskoff and E. Vanden-Eijnden · 2018
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
M.S. Advani, A.M. Saxe, and H. Sompolinsky · 2020
Later among the works it cites.
Double trouble in double descent : Bias and variance(s) in the lazy regime
S. d’Ascoli, M. Refinetti, G. Biroli, and F. Krzakala · 2020
Later among the works it cites.
Understanding double descent requires a fine-grained bias-variance decomposition
B. Adlam and J. Pennington · 2020
Later among the works it cites.
Scaling description of generalization with number of parameters in deep learning
M. Geiger, A. Jacot, S. Spigler, F. Gabriel, L. Sagun, S. d’Ascoli, G. Biroli, C. Hongler, and M. Wyart · 2020
Later among the works it cites.
Benign overfitting in linear regression
P.L. Bartlett, P.M. Long, G. Lugosi, and A. Tsigler · 2020
Later among the works it cites.
Bad global minima exist and sgd can reach them
S. Liu, D. Papailiopoulos, and D. Achlioptas · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the global convergence of gradient descent for over-parameterized models using optimal transport
L. Chizat and F. Bach · 2018
Cited alongside, same era.
Mixed precision training
P. Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, and H. Wu · 2018
Cited alongside, same era.
A jamming transition from under-to over-parametrization affects generalization in deep learning
S Spigler, M Geiger, S d’Ascoli, L Sagun, G Biroli, and M Wyart · 2019
Cited alongside, same era.
The implicit bias of gradient descent on nonseparable data
Z. Ji and M. Telgarsky · 2019
Cited alongside, same era.
Implicit regularization in deep matrix factorization
S. Arora, N. Cohen, W. Hu, and Y. Luo · 2019
Cited alongside, same era.
The generalization error of random features regression: Precise asymptotics and the double descent curve
S. Mei and A. Montanari · 2019
Cited alongside, same era.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
M. Belkin, D. Hsu, S. Ma, and S. Mandal · 2019
Cited alongside, same era.
Later among the works it cites.
A constructive prediction of the generalization error across scales
J.S. Rosenfeld, A. Rosenfeld, Y. Belinkov, and N. Shavit · 2020
Later among the works it cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T.B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Later among the works it cites.
Scaling laws for autoregressive generative modeling
T. Henighan, J. Kaplan, M. Katz, M. Chen, C. Hesse, J. Jackson, H. Jun, T.B. Brown, P. Dhariwal, S. Gray, et al · 2020
Later among the works it cites.
Disentangling feature and lazy training in deep neural networks
M. Geiger, S. Spigler, A. Jacot, and M. Wyart · 2020
Later among the works it cites.
Landscape connectivity and dropout stability of sgd solutions for over-parameterized neural networks
A. Shevchenko and M. Mondelli · 2020
Later among the works it cites.
Rethinking bias-variance trade-off for generalization of neural networks
Zitong Yang, Yaodong Yu, Chong You, Jacob Steinhardt, and Yi Ma · 2020
Later among the works it cites.
What causes the test error? going beyond bias-variance via anova
L. Lin and E. Dobriban · 2021
Closest in time.
Explaining neural scaling laws
Y. Bahri, E. Dyer, J. Kaplan, J. Lee, and U. Sharma · 2021
Closest in time.
Classifying high-dimensional gaussian mixtures: Where kernel methods fail and neural networks succeed
M. Refinetti, S. Goldt, F. Krzakala, and L. Zdeborova · 2021
Closest in time.
On connectivity of solutions in deep learning: The role of over-parameterization and feature quality
Q. Nguyen, P. Brechet, and M. Mondelli · 2021
Closest in time.
Revisiting resnets: Improved training and scaling strategies
I. Bello, W. Fedus, X. Du, E. D. Cubuk, A. Srinivas, T.-Y. Lin, J. Shlens, and B. Zoph · 2021
Closest in time.
Surprises in high-dimensional ridgeless least squares interpolation
T. Hastie, A. Montanari, S. Rosset, and R.J. Tibshirani · 2022
Closest in time.