Fetching the paper…
Reading the bibliography…
We consider an underdetermined noisy linear regression model where the minimum-norm interpolating predictor is known to be consistent, and ask: can uniform convergence in a norm ball, or at least (following Nagarajan and Kolter) the subset of a norm ball that the algorithm selects on a typical input set, explain this success? We show that uniformly bounding the difference between empirical and population errors cannot show any learning in the norm ball, and cannot show consistency for any set, even one depending on the exact algorithm and distribution.
“Uniform convergence may be unable to explain generalization in deep learning”
Vaishnavh Nagarajan and J. Kolter · 1902
Earlier work this paper cites.
“Two models of double descent for weak features”
Mikhail Belkin, Daniel Hsu and Ji Xu · 1903
Earlier work this paper cites.
“Surprises in High-Dimensional Ridgeless Least Squares Interpolation”, 2019
Trevor Hastie, Andrea Montanari, Saharon Rosset and Ryan. Tibshirani · 1903
Earlier work this paper cites.
“Harmless interpolation of noisy data in regression”
Vidya Muthukumar, Kailas Vodrahalli, Vignesh Subramanian and Anant Sahai · 1903
Earlier work this paper cites.
“Benign overfitting in linear regression”
Peter. Bartlett, Philip. Long, Gábor Lugosi and Alexander Tsigler · 1906
Earlier work this paper cites.
Song Mei and Andrea Montanari · 1908
Earlier work this paper cites.
Andrea Montanari, Feng Ruan, Youngtak Sohn and Jun Yan · 1911
Earlier work this paper cites.
“Exact expressions for double descent and implicit regularization via surrogate random design”
Michał Dereziński, Feynman Liang and Michael. Mahoney · 1912
Earlier work this paper cites.
“Deep Double Descent: Where Bigger Models and More Data Hurt”
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak and Ilya Sutskever · 1912
Earlier work this paper cites.
Jeffrey Negrea, Gintare Dziugaite and Daniel. Roy · 1912
Earlier work this paper cites.
“A theory of the learnable”
Leslie Valiant · 1984
Earlier work this paper cites.
“Regression Shrinkage and Selection via the Lasso”
Robert Tibshirani · 1996
Earlier work this paper cites.
“Overfitting Can Be Harmless for Basis Pursuit: Only to a Degree”
Peizhong Ju, Xiaojun Lin and Jia Liu · 2002
Cited alongside, same era.
“Some Extensions of an Inequality of Vapnik and Chervonenkis”
Dmitriy Panchenko · 2002
Cited alongside, same era.
“Convex Optimization”
Stephen Boyd and Lieven Vandenberghe · 2004
Cited alongside, same era.
“The Dantzig selector: Statistical estimation when p p is much larger than n n ”
Emmanuel Candes and Terence Tao · 2005
Cited alongside, same era.
“On Model Selection Consistency of Lasso”
Peng Zhao and Bin Yu · 2006
Cited alongside, same era.
“Optimal rates for regularized least-squares algorithm”
Andrea Caponnetto and Ernesto De Vito · 2007
Cited alongside, same era.
“High-dimensional dynamics of generalization error in neural networks”, 2017
Madhu Advani and Andrew Saxe · 2017
Later among the works it cites.
“Concentration Inequalities and Moment Bounds for Sample Covariance Operators”
Vladimir Koltchinskii and Karim Lounici · 2017
Later among the works it cites.
“Understanding deep learning requires rethinking generalization”
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht and Oriol Vinyals · 2017
Later among the works it cites.
Mikhail Belkin, Daniel. Hsu and Partha Mitra · 2018
Later among the works it cites.
“To understand deep learning we need to understand kernel learning”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“On the complexity of linear prediction: Risk bounds, margin bounds, and regularization”
Sham Kakade, Karthik Sridharan and Ambuj Tewari · 2009
Cited alongside, same era.
“Optimistic Rates for Learning with a Smooth Loss”, 2010
Nathan Srebro, Karthik Sridharan and Ambuj Tewari · 2010
Cited alongside, same era.
“Foundations of Machine Learning”
Mehryar Mohri, Afshin Rostamizadeh and Ameet Talwalkar · 2012
Cited alongside, same era.
“Assumptionless consistency of the Lasso”, 2013
Sourav Chatterjee · 2013
Cited alongside, same era.
“Understanding Machine Learning: From Theory to Algorithms”
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
“In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning”
Behnam Neyshabur, Ryota Tomioka and Nathan Srebro · 2015
Cited alongside, same era.
Mikhail Belkin, Siyuan Ma and Soumik Mandal · 2018
Later among the works it cites.
“A jamming transition from under- to over-parametrization affects generalization in deep learning”
Stefano Spigler, Mario Geiger, Stéphane d’Ascoli, Levent Sagun, Giulio Biroli and Matthieu Wyart · 2018
Later among the works it cites.
“Reconciling modern machine learning practice and the bias-variance trade-off”
Mikhail Belkin, Daniel Hsu, Siyuan Ma and Soumik Mandal · 2019
Later among the works it cites.
“Does data interpolation contradict statistical optimality?”
Mikhail Belkin, Alexander Rakhlin and Alexandre. Tsybakov · 2019
Later among the works it cites.
“High-Dimensional Statistics: A Non-Asymptotic Viewpoint”
Martin. Wainwright · 2019
Later among the works it cites.
“Generalization of Two-layer Neural Networks: An Asymptotic Viewpoint”
Jimmy Ba, Murat Erdogdu, Taiji Suzuki, Denny Wu and Tianzong Zhang · 2020
Closest in time.
“Minimax rates of estimation for high-dimensional linear regression over l q l_{q} -balls”
Garvesh Raskutti, Martin. Wainwright and Bin Yu · 2042
Closest in time.