Fetching the paper…
Reading the bibliography…
Modern deep learning models with great expressive power can be trained to overfit the training data but still generalize well.
Feature purification: How adversarial training performs robust deep learning
Allen-Zhu, Z. and Li, Y · 2005
Earlier work this paper cites.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Allen-Zhu, Z. and Li, Y · 2012
Earlier work this paper cites.
Lin, M., Chen, Q., and Yan, S · 2013
Earlier work this paper cites.
Train faster, generalize better: stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y · 2016
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
Bartlett, P. L., Foster, D. J., and Telgarsky, M. J · 2017
Earlier work this paper cites.
Generalization bounds of sgld for non-convex learning: Two theoretical viewpoints
Mou, W., Wang, L., Zhai, X., and Zheng, K · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2017
Earlier work this paper cites.
To understand deep learning we need to understand kernel learning
Belkin, M., Ma, S., and Mandal, S · 2018
Earlier work this paper cites.
Stability and convergence trade-off of iterative optimization algorithms
Chen, Y., Jin, C., and Yu, B · 2018
Earlier work this paper cites.
The total variation distance between high-dimensional gaussians
Devroye, L., Mehrabian, A., and Reddad, T · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
High-Dimensional Probability: An Introduction with Applications in Data Science
Vershynin, R · 2018
Earlier work this paper cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Belkin, M., Hsu, D., Ma, S., and Mandal, S · 2019
Cited alongside, same era.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Li, Y., Wei, C., and Ma, T · 2019
Cited alongside, same era.
The neural tangent kernel in high dimensions: Triple descent and a multi-scale theory of generalization
Adlam, B. and Pennington, J · 2020
Cited alongside, same era.
Benign overfitting in linear regression
Bartlett, P. L., Long, P. M., Lugosi, G., and Tsigler, A · 2020
Cited alongside, same era.
Two models of double descent for weak features
Belkin, M., Hsu, D., and Xu, J · 2020
Cited alongside, same era.
Just interpolate: Kernel “ridgeless” regression can generalize
Liang, T. and Rakhlin, A · 2020
Towards an understanding of benign overfitting in neural networks
Li, Z., Zhou, Z.-H., and Gretton, A · 2021
Later among the works it cites.
Classification vs regression in overparameterized regimes: Does the loss function matter?
Muthukumar, V., Narang, A., Subramanian, V., Belkin, M., Hsu, D., and Sahai, A · 2021
Later among the works it cites.
Benign overfitting in binary classification of gaussian mixtures
Wang, K. and Thrampoulidis, C · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2021
Later among the works it cites.
Benign overfitting in two-layer convolutional neural networks
Cao, Y., Chen, Z., Belkin, M., and Gu, Q · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the multiple descent of minimum-norm interpolants and restricted lower isometry of kernels
Liang, T., Rakhlin, A., and Zhai, X · 2020
Cited alongside, same era.
On the optimal weighted ℓ 2 \ell_{2} regularization in overparameterized linear regression
Wu, D. and Xu, J · 2020
Cited alongside, same era.
Risk bounds for over-parameterized maximum margin classification on sub-gaussian mixtures
Cao, Y., Gu, Q., and Belkin, M · 2021
Cited alongside, same era.
Finite-sample analysis of interpolating linear classifiers in the overparameterized regime
Chatterji, N. S. and Long, P. M · 2021
Cited alongside, same era.
Proxy convexity: A unified framework for the analysis of neural networks trained by gradient descent
Frei, S. and Gu, Q · 2021
Cited alongside, same era.
Provable generalization of sgd-trained neural networks of any width in the presence of adversarial label noise
Frei, S., Cao, Y., and Gu, Q · 2021
Cited alongside, same era.
Chatterji, N. S. and Long, P. M · 2022
Later among the works it cites.
Benign overfitting without linearity: Neural network classifiers trained by gradient descent for noisy linear data
Frei, S., Chatterji, N. S., and Bartlett, P · 2022
Later among the works it cites.
Surprises in high-dimensional ridgeless least squares interpolation
Hastie, T., Montanari, A., Rosset, S., and Tibshirani, R. J · 2022
Later among the works it cites.
The interpolation phase transition in neural networks: Memorization and generalization under lazy training
Montanari, A. and Zhong, Y · 2022
Later among the works it cites.
The implicit bias of benign overfitting
Shamir, O · 2022
Later among the works it cites.
Data augmentation as feature manipulation: a story of desert cows and grass cows
Shen, R. and Bubeck, S · 2022
Later among the works it cites.
Benign overfitting of constant-stepsize sgd for linear regression
Zou, D., Wu, J., Braverman, V., Gu, Q., and Kakade, S · 2022
Later among the works it cites.