Fetching the paper…
Reading the bibliography…
The practical success of overparameterized neural networks has motivated the recent scientific study of interpolating methods, which perfectly fit their training data.
Two models of double descent for weak features
Mikhail Belkin, Daniel Hsu, and Ji Xu · 1903
Earlier work this paper cites.
Nearest neighbor pattern classification
Thomas Cover and Peter Hart · 1967
Earlier work this paper cites.
The hilbert kernel regression estimate
Luc Devroye, Laszlo Györfi, and Adam Krzyżak · 1998
Earlier work this paper cites.
The mnist database of handwritten digits
Yann LeCun · 1998
Earlier work this paper cites.
Consistency and robustness of kernel-based regression in convex risk minimization
Andreas Christmann and Ingo Steinwart · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky et al · 2009
Earlier work this paper cites.
MNIST handwritten digit database
Yann LeCun and Corinna Cortes · 2010
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Rates of convergence for nearest neighbor classification
Kamalika Chaudhuri and Sanjoy Dasgupta · 2014
Earlier work this paper cites.
Xsede: Accelerating scientific discovery
J. Towns, T. Cockerill, M. Dahan, I. Foster, K. Gaither, A. Grimshaw, V. Hazlewood, S. Lathrop, D. Lifka, G. D. Peterson, R. Roskies, J. Scott, and N. Wilkins-Diehr · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Statistical physics of inference: Thresholds and algorithms
Lenka Zdeborová and Florent Krzakala · 2016
Earlier work this paper cites.
High-dimensional dynamics of generalization error in neural networks
Madhu S Advani and Andrew M Saxe · 2017
Earlier work this paper cites.
Diving into the shallows: a computational perspective on large-scale shallow learning
Siyuan Ma and Mikhail Belkin · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
Overfitting or perfect fitting? risk bounds for classification and regression rules that interpolate
Mikhail Belkin, Daniel J. Hsu, and Partha Mitra · 2018
Earlier work this paper cites.
To understand deep learning we need to understand kernel learning
Mikhail Belkin, Siyuan Ma, and Soumik Mandal · 2018
Earlier work this paper cites.
Cinic-10 is not imagenet or cifar-10
Luke Nicholas Darlow, Elliot J. Crowley, Antreas Antoniou, and Amos J. Storkey · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Just interpolate: Kernel" ridgeless" regression can generalize
Tengyuan Liang and Alexander Rakhlin · 2018
Cited alongside, same era.
A continuous-time view of early stopping for least squares regression
Alnur Ali, J. Zico Kolter, and Ryan J. Tibshirani · 2019
Cited alongside, same era.
Does data interpolation contradict statistical optimality?
Mikhail Belkin, Alexander Rakhlin, and Alexandre B. Tsybakov · 2019
Cited alongside, same era.
Distributional generalization: A new kind of generalization
Preetum Nakkiran and Yamini Bansal · 2020
Later among the works it cites.
Asymptotic learning curves of kernel methods: empirical data versus teacher–student paradigm
Stefano Spigler, Mario Geiger, and Matthieu Wyart · 2020
Later among the works it cites.
Benign overfitting in ridge regression
Alexander Tsigler and Peter L Bartlett · 2020
Later among the works it cites.
Explaining neural scaling laws
Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma · 2021
Later among the works it cites.
Deep learning: a statistical viewpoint
Peter L Bartlett, Andrea Montanari, and Alexander Rakhlin · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dynamics of stochastic gradient descent for two-layer neural networks in the teacher-student setup
Sebastian Goldt, Madhu Advani, Andrew M Saxe, Florent Krzakala, and Lenka Zdeborová · 2019
Cited alongside, same era.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2019
Cited alongside, same era.
The generalization error of random features regression: Precise asymptotics and double descent curve
Song Mei and Andrea Montanari · 2019
Cited alongside, same era.
Consistency of interpolation with laplace kernels is a high-dimensional phenomenon
Alexander Rakhlin and Xiyu Zhai · 2019
Cited alongside, same era.
Statistical mechanics of deep learning
Yasaman Bahri, Jonathan Kadmon, Jeffrey Pennington, Sam S Schoenholz, Jascha Sohl-Dickstein, and Surya Ganguli · 2020
Cited alongside, same era.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2020
Cited alongside, same era.
Deep equals shallow for relu networks in kernel regimes
Alberto Bietti and Francis Bach · 2020
Cited alongside, same era.
Later among the works it cites.
Spectral bias and task-model alignment explain generalization in kernel regression and infinitely wide neural networks
Abdulkadir Canatar, Blake Bordelon, and Cengiz Pehlevan · 2021
Later among the works it cites.
Foolish crowds support benign overfitting
Niladri S Chatterji and Philip M Long · 2021
Later among the works it cites.
The interplay between implicit bias and benign overfitting in two-layer linear networks
Niladri S Chatterji, Philip M Long, and Peter L Bartlett · 2021
Later among the works it cites.
Generalization error rates in kernel regression: The crossover from the noiseless to noisy regime
Hugo Cui, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová · 2021
Later among the works it cites.
On the universality of the double descent peak in ridgeless regression
David Holzmüller · 2021
Later among the works it cites.
Early-stopped neural networks are consistent
Ziwei Ji, Justin Li, and Matus Telgarsky · 2021
Later among the works it cites.
Uniform convergence of interpolators: Gaussian width, norm bounds and benign overfitting
Frederic Koehler, Lijia Zhou, Danica J Sutherland, and Nathan Srebro · 2021
Later among the works it cites.
James B Simon, Madeline Dickens, Dhruva Karkada, and Michael R DeWeese · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Later among the works it cites.
Benign overfitting in two-layer convolutional neural networks
Yuan Cao, Zixiang Chen, Mikhail Belkin, and Quanquan Gu · 2022
Closest in time.
Spencer Frei, Niladri S Chatterji, and Peter L Bartlett · 2022
Closest in time.
Wide and deep neural networks achieve optimality for classification, 2022
Adityanarayanan Radhakrishnan, Mikhail Belkin, and Caroline Uhler · 2022
Closest in time.
Reverse engineering the neural tangent kernel
James B. Simon, Sajant Anand, and Michael Robert DeWeese · 2022
Closest in time.