Fetching the paper…
Reading the bibliography…
Recent empirical and theoretical studies have established the generalization capabilities of large machine learning models that are trained to (approximately or exactly) fit noisy data.
Benign overfitting in linear regression
Bartlett, P. L., Long, P. M., Lugosi, G., and Tsigler, A. (2019) · 1906
Earlier work this paper cites.
Adversarial training can hurt generalization
Raghunathan, A., Xie, S. M., Yang, F., Duchi, J. C., and Liang, P. (2019) · 1906
Earlier work this paper cites.
Adversarial classification
Dalvi, N., Domingos, P., Sanghai, S., and Verma, D. (2004) · 2004
Earlier work this paper cites.
How benign is benign overfitting?
Sanyal, A., Dokania, P. K., Kanade, V., and Torr, P. H. (2020) · 2007
Earlier work this paper cites.
Methodologies in spectral analysis of large dimensional random matrices, a review
Bai, Z. D. (2008) · 2008
Earlier work this paper cites.
The spectrum of kernel random matrices
El Karoui, N. (2010) · 2010
Earlier work this paper cites.
Evasion attacks against machine learning at test time
Biggio, B., Corona, I., Maiorca, D., Nelson, B., Šrndić, N., Laskov, P., Giacinto, G., and Roli, F. (2013) · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R. (2013) · 2013
Earlier work this paper cites.
Margins, shrinkage, and boosting
Telgarsky, M. (2013) · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Goodfellow, I. J., Shlens, J., and Szegedy, C. (2014) · 2014
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P. (2016) · 2016
Earlier work this paper cites.
Concentration inequalities and moment bounds for sample covariance operators
Koltchinskii, V. and Lounici, K. (2017) · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A. (2017) · 2017
Earlier work this paper cites.
Exploring generalization in deep learning
Neyshabur, B., Bhojanapalli, S., McAllester, D., and Srebro, N. (2017) · 2017
Earlier work this paper cites.
The marginal value of adaptive gradient methods in machine learning
Wilson, A. C., Roelofs, R., Stern, M., Srebro, N., and Recht, B. (2017) · 2017
Earlier work this paper cites.
Explaining the success of adaboost and random forests as interpolating classifiers
Wyner, A. J., Olson, M., Bleich, J., and Mease, D. d. (2017) · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O. (2017) · 2017
Earlier work this paper cites.
To understand deep learning we need to understand kernel learning
Belkin, M., Ma, S., and Mandal, S. (2018) · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C. (2018) · 2018
Earlier work this paper cites.
Adversarially robust generalization requires more data
Schmidt, L., Santurkar, S., Tsipras, D., Talwar, K., and Madry, A. (2018) · 2018
Cited alongside, same era.
Are adversarial examples inevitable?
Shafahi, A., Huang, W. R., Studer, C., Feizi, S., and Goldstein, T. (2018) · 2018
Cited alongside, same era.
Robustness may be at odds with accuracy
Tsipras, D., Santurkar, S., Engstrom, L., Turner, A., and Madry, A. (2018) · 2018
Cited alongside, same era.
High-dimensional probability: An introduction with applications in data science
Vershynin, R. (2018) · 2018
Cited alongside, same era.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Belkin, M., Hsu, D., Ma, S., and Mandal, S. (2019) · 2019
Cited alongside, same era.
Interpolation can hurt robust generalization even when there is no noise
Donhauser, K., Tifrea, A., Aerni, M., Heckel, R., and Yang, F. (2021) · 2021
Later among the works it cites.
Exploring architectural ingredients of adversarially robust deep neural networks
Huang, H., Wang, Y., Erfani, S., Gu, Q., Bailey, J., and Ma, X. (2021) · 2021
Later among the works it cites.
Towards an understanding of benign overfitting in neural networks
Li, Z., Zhou, Z.-H., and Gretton, A. (2021) · 2021
Later among the works it cites.
Classification vs regression in overparameterized regimes: Does the loss function matter?
Muthukumar, V., Narang, A., Subramanian, V., Belkin, M., Hsu, D., and Sahai, A. (2021) · 2021
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I. (2021) · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Cao, Y. and Gu, Q. (2019) · 2019
Cited alongside, same era.
Convergence of adversarial training in overparametrized neural networks
Gao, R., Cai, T., Li, H., Hsieh, C.-J., Wang, L., and Lee, J. D. (2019) · 2019
Cited alongside, same era.
Adversarial examples are not bugs, they are features
Ilyas, A., Santurkar, S., Tsipras, D., Engstrom, L., Tran, B., and Madry, A. (2019) · 2019
Cited alongside, same era.
Improving adversarial robustness requires revisiting misclassified examples
Wang, Y., Zou, D., Yi, J., Bailey, J., Ma, X., and Gu, Q. (2019) · 2019
Cited alongside, same era.
Theoretically principled trade-off between robustness and accuracy
Zhang, H., Yu, Y., Jiao, J., Xing, E., El Ghaoui, L., and Jordan, M. (2019) · 2019
Cited alongside, same era.
The neural tangent kernel in high dimensions: Triple descent and a multi-scale theory of generalization
Adlam, B. and Pennington, J. (2020) · 2020
Cited alongside, same era.
Two models of double descent for weak features
Belkin, M., Hsu, D., and Xu, J. (2020) · 2020
Cited alongside, same era.
Do wider neural networks really help adversarial robustness?
Wu, B., Chen, J., Cai, D., He, X., and Gu, Q. (2021) · 2021
Later among the works it cites.
Provable robustness of adversarial training for learning halfspaces with noise
Zou, D., Frei, S., and Gu, Q. (2021) · 2021
Later among the works it cites.
Hassani, H. and Javanmard, A. (2022) · 2022
Later among the works it cites.
Surprises in high-dimensional ridgeless least squares interpolation
Hastie, T., Montanari, A., Rosset, S., and Tibshirani, R. J. (2022) · 2022
Later among the works it cites.
The implicit bias of benign overfitting
Shamir, O. (2022) · 2022
Later among the works it cites.
Binary classification of gaussian mixtures: Abundance of support vectors, benign overfitting, and regularization
Wang, K. and Thrampoulidis, C. (2022) · 2022
Later among the works it cites.
A universal law of robustness via isoperimetry
Bubeck, S. and Sellke, M. (2023) · 2023
Later among the works it cites.
Benign overfitting in adversarially robust linear classification
Chen, J., Cao, Y., and Gu, Q. (2023) · 2023
Later among the works it cites.
Provable tradeoffs in adversarially robust classification
Dobriban, E., Hassani, H., Hong, D., and Robey, A. (2023) · 2023
Later among the works it cites.
Simon, J. B., Karkada, D., Ghosh, N., and Belkin, M. (2023) · 2023
Later among the works it cites.
Benign overfitting in ridge regression
Tsigler, A. and Bartlett, P. L. (2023) · 2023
Later among the works it cites.
Benign overfitting in multiclass classification: All roads lead to interpolation
Wang, K., Muthukumar, V., and Thrampoulidis, C. (2023) · 2023
Later among the works it cites.
Benign overfitting in deep neural networks under lazy training
Zhu, Z., Liu, F., Chrysos, G., Locatello, F., and Cevher, V. (2023) · 2023
Later among the works it cites.
Precise tradeoffs in adversarial training for linear regression
Javanmard, A., Soltanolkotabi, M., and Hassani, H. (2020) · 2078
Closest in time.