Fetching the paper…
Reading the bibliography…
The recent success of neural network models has shone light on a rather surprising statistical phenomenon: statistical models that perfectly fit noisy data can generalize well to unseen test data.
“Introduction to algorithms”
Thomas Cormen, Charles Leiserson, Ronald Rivest and Clifford Stein · 2009
Earlier work this paper cites.
“Smallest singular value of a random rectangular matrix”
Mark Rudelson and Roman Vershynin · 2009
Earlier work this paper cites.
“Non-asymptotic theory of random matrices: extreme singular values”
Mark Rudelson and Roman Vershynin · 2010
Earlier work this paper cites.
“Hanson-Wright inequality and sub-Gaussian concentration”
Mark Rudelson and Roman Vershynin · 2013
Earlier work this paper cites.
“In search of the real inductive bias: on the role of implicit regularization in deep learning.”
Behnam Neyshabur, Ryota Tomioka and Nathan Srebro · 2015
Earlier work this paper cites.
“Small ball probabilities for linear images of high-dimensional distributions”
Mark Rudelson and Roman Vershynin · 2015
Earlier work this paper cites.
“Implicit regularization in matrix factorization”
Suriya Gunasekar, Blake Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur and Nathna Srebro · 2017
Earlier work this paper cites.
“Concentration inequalities and moment bounds for sample covariance operators”
Vladimir Koltchinskii and Karim Lounici · 2017
Earlier work this paper cites.
“Understanding deep learning requires rethinking generalization”
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht and Oriol Vinyals · 2017
Earlier work this paper cites.
“Overfitting or perfect fitting? Risk bounds for classification and regression rules that interpolate”
Mikhail Belkin, Daniel Hsu and Partha Mitra · 2018
Earlier work this paper cites.
“Characterizing implicit bias in terms of optimization geometry”
Suriya Gunasekar, Jason Lee, Daniel Soudry and Nathan Srebro · 2018
Earlier work this paper cites.
“Implicit bias of gradient descent on linear convolutional networks”
Suriya Gunasekar, Jason Lee, Daniel Soudry and Nathan Srebro · 2018
Earlier work this paper cites.
“The implicit bias of gradient descent on separable data”
Daniel Soudry, Elad Hoffer, Mor Nacson, Suriya Gunasekar and Nathan Srebro · 2018
Earlier work this paper cites.
“High-dimensional probability: An introduction with applications in data science”
Roman Vershynin · 2018
Earlier work this paper cites.
“Implicit regularization in deep matrix factorization”
Sanjeev Arora, Nadav Cohen, Wei Hu and Yuping Luo · 2019
Earlier work this paper cites.
“Reconciling modern machine-learning practice and the classical bias–variance trade-off”
Mikhail Belkin, Daniel Hsu, Siyuan Ma and Soumik Mandal · 2019
Earlier work this paper cites.
“The implicit bias of gradient descent on nonseparable data”
Ziwei Ji and Matus Telgarsky · 2019
Earlier work this paper cites.
“The generalization error of random features regression: Precise asymptotics and the double descent curve”
Song Mei and Andrea Montanari · 2019
Cited alongside, same era.
Andrea Montanari, Feng Ruan, Youngtak Sohn and Jun Yan · 2019
Cited alongside, same era.
“Stochastic gradient descent on separable data: Exact convergence with a fixed learning rate”
Mor Nacson, Nathan Srebro and Daniel Soudry · 2019
Cited alongside, same era.
“Benign overfitting in linear regression”
Peter Bartlett, Philip Long, Gábor Lugosi and Alexander Tsigler · 2020
Cited alongside, same era.
“Interpolation under latent factor regression models”
Florentina Bunea, Seth Strimas-Mackey and Marten Wegkamp · 2020
Cited alongside, same era.
“Kernel and rich regimes in overparametrized models”
Blake Woodworth, Suriya Gunasekar, Jason Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry and Nathan Srebro · 2020
Later among the works it cites.
“On the optimal weighted ℓ 2 \ell_{2} regularization in overparameterized linear regression”
Denny Wu and Ji Xu · 2020
Later among the works it cites.
“Rethinking bias-variance trade-off for generalization of neural networks”
Zitong Yang, Yaodong Yu, Chong You, Jacob Steinhardt and Yi Ma · 2020
Later among the works it cites.
“On the implicit bias of initialization shape: beyond infinitesimal mirror descent”
Shahar Azulay, Edward Moroshko, Mor Nacson, Blake. Woodworth, Nathan Srebro, Amir Globerson and Daniel Soudry · 2021
Closest in time.
“Deep learning: a statistical viewpoint”
Peter Bartlett, Andrea Montanari and Alexander Rakhlin · 2021
Closest in time.
“Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“On the robustness of the minimum ℓ 2 \ell_{2} interpolator”
Geoffrey Chinot and Matthieu Lerasle · 2020
Cited alongside, same era.
“On the robustness of minimum-norm interpolators”
Geoffrey Chinot, Matthias Löffler and Sara van Geer · 2020
Cited alongside, same era.
“The implicit bias of depth: how incremental learning drives generalization”
Daniel Gissin, Shai Shalev-Shwartz and Amit Daniely · 2020
Cited alongside, same era.
“The optimal ridge penalty for real-world high-dimensional data can be zero or negative due to the implicit ridge regularization”
Dmitry Kobak, Jonathan Lomond and Benoit Sanchez · 2020
Cited alongside, same era.
“Just interpolate: Kernel “ridgeless" regression can generalize”
Tengyuan Liang and Alexander Rakhlin · 2020
Cited alongside, same era.
“On the multiple descent of minimum-norm interpolants and restricted lower isometry of kernels”
Tengyuan Liang, Alexander Rakhlin and Xiyu Zhai · 2020
Cited alongside, same era.
Tengyuan Liang and Pragya Sur · 2020
Cited alongside, same era.
Mikhail Belkin · 2021
Closest in time.
“Finite-sample analysis of interpolating linear classifiers in the overparameterized regime”
Niladri Chatterji and Philip Long · 2021
Closest in time.
“On the proliferation of support vectors in high dimensions”
Daniel Hsu, Vidya Muthukumar and Ji Xu · 2021
Closest in time.
“Inductive Bias of Multi-Channel Linear Convolutional Networks with Bounded Weight Norm”
Meena Jagadeesan, Ilya Razenshteyn and Suriya Gunasekar · 2021
Closest in time.
“Uniform convergence of interpolators: Gaussian width, norm bounds and benign overfitting”
Frederic Koehler, Lijia Zhou, Danica Sutherland and Nathan Srebro · 2021
Closest in time.
“Towards an understanding of benign overfitting in neural networks”
Zhu Li, Zhi-Hua Zhou and Arthur Gretton · 2021
Closest in time.
“Classification vs regression in overparameterized regimes: Does the loss function matter?”
Vidya Muthukumar, Adhyyan Narang, Vignesh Subramanian, Mikhail Belkin, Daniel Hsu and Anant Sahai · 2021
Closest in time.
“Benign overfitting in binary classification of gaussian mixtures”
Ke Wang and Christos Thrampoulidis · 2021
Closest in time.
“A unifying view on implicit bias in training linear neural networks”
Chulhee Yun, Shankar Krishnan and Hossein Mobahi · 2021
Closest in time.
“A model of double descent for high-dimensional binary linear classification”
Zeyu Deng, Abla Kammoun and Christos Thrampoulidis · 2022
Closest in time.
“Surprises in high-dimensional ridgeless least squares interpolation”
Trevor Hastie, Andrea Montanari, Saharon Rosset and Ryan Tibshirani · 2022
Closest in time.