Fetching the paper…
Reading the bibliography…
Double-descent curves in neural networks describe the phenomenon that the generalisation error initially descends with increasing parameters, then grows after reaching an optimal number of parameters which is less than the number of data points, but then descends again in the overparameterized regime.
“Functions of positive and negative type, and their connection the theory of integral equations”
J. Mercer · 1909
Earlier work this paper cites.
“On the reciprocal of the general algebraic matrix”
E.H. Moore · 1920
Earlier work this paper cites.
“Characteristic vectors of bordered matrices with infinite dimensions”
E. Wigner · 1955
Earlier work this paper cites.
“Distribution of eigenvalues for some sets of random matrices”
V.A. Marchenko and L.A. Pastur · 1967
Earlier work this paper cites.
“The Numerical Treatment of Integral Equations”
C… Baker · 1977
Earlier work this paper cites.
“Linear and Nonlinear Extension of the Pseudo-Inverse Solution for Learning Boolean Functions”
F. Vallet, J.-G. Cailton and P. Refregier · 1989
Earlier work this paper cites.
“Statistical mechanics of learning from examples”
H.. Seung and H. Sompolinsky · 1992
Earlier work this paper cites.
“Bayesian Learning for Neural Networks”, 1994
R. Neal · 1994
Earlier work this paper cites.
“The Nature of Statistical Learning Theory”
V. Vapnik · 1995
Earlier work this paper cites.
“Necessary and Sufficient Condition that the Limit of Stieltjes Transforms is a Stieltjes Transform”
J.S. Geronimo and T.P. Hill · 2002
Earlier work this paper cites.
“Learning Curves for Gaussian Process Regression: Approximations and Bounds”
Peter Sollich and Anason Halees · 2002
Earlier work this paper cites.
“Measures, Integrals and Martingales”
R.. Schilling · 2005
Earlier work this paper cites.
“Mercer’s Theorem, Feature Maps, and Smoothing”
H.Q. Minh, P. Niyogi and Y. Yao · 2006
Earlier work this paper cites.
“Gaussian Processes for Machine Learning”
C.E. Rasmussen and C.K.I. Williams · 2006
Earlier work this paper cites.
“The spectrum of kernel random matrices”
Noureddine Karoui · 2010
Earlier work this paper cites.
“Statistical mechanics of learning”
A. Engel, Germany Otto-von Guericke-Universität and C. den Broeck · 2012
Earlier work this paper cites.
“Imagenet classification with deepconvolutional neural networks.”
A. Krizhevsky, I. Sutskever and G.. Hinton · 2012
Earlier work this paper cites.
“The MNIST database”
Yann LeCun · 2012
Earlier work this paper cites.
“Foundations of Machine Learning”
M. Mohri, A. Rostamizadeh and A.Talwalkar · 2012
Earlier work this paper cites.
“THE SPECTRUM OF RANDOM INNER-PRODUCT KERNEL MATRICES”
XIUYUAN CHENG and AMIT SINGER · 2013
Earlier work this paper cites.
“Almost-commuting matrices are almost jointly diagonalizable”
K. Glashoff and M.. Bronstein · 2013
Earlier work this paper cites.
“Deep speech: Scaling up end-to-end speech recognition.”
A. Hannun et al · 2014
Earlier work this paper cites.
“The arrival of the frequent: how bias in genotype-phenotype maps can steer populations to local optima”
S. Schaper and A.. Louis · 2014
Earlier work this paper cites.
“Understanding machine learning: From theory to algorithms”
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
“Very deep convolutional networks for large-scale image recognition”
K. Simonyan and A Zisserman · 2014
Earlier work this paper cites.
“On the limiting spectral distribution for a large class of symmetric random matrices with correlated entries”
M. Banna, F. Merlevede and M. Peligrad · 2015
Cited alongside, same era.
“The spectral norm of random inner-product kernel matrices”
Zhou Fan and Andrea Montanari · 2015
Cited alongside, same era.
“Deep learning”
Y. LeCun, Y. Bengio and G. Hinton · 2015
Cited alongside, same era.
“Deep learning in neural networks: An overview”
J\"urgen Schmidhuber · 2015
Cited alongside, same era.
“Toward Deeper Understanding of Neural Networks: The Power of Initialization and a Dual View on Expressivity”
A. Daniely, R. Frostig and Y. Singer · 2016
Cited alongside, same era.
“Exponential expressivity in deep neural networks through transient chaos”
B. Poole et al · 2016
Cited alongside, same era.
“Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural Networks”
B. Bordelon, A. Canatar and C. Pehlevan · 2020
Later among the works it cites.
“Towards Understanding the Spectral Bias of Deep Learning”
Y. Cao et al · 2020
Later among the works it cites.
“Generalization Error of Generalized Linear Models in High Dimensions”
Melikasadat Emami et al · 2020
Later among the works it cites.
“Spectra of the Conjugate Kernel and Neural Tangent Kernel for Linear-Width Neural Networks”
Z. Fan and Z. Wang · 2020
Later among the works it cites.
M. Geiger, L. Petrini and M Wyart · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Mastering the game of Go with deep neural networks and tree search.”
D. Silver et al · 2016
Cited alongside, same era.
“Input–output maps are strongly biased towards simple outputs”
K. Dingle, C.. Camargo and A.. Louis · 2018
Cited alongside, same era.
“Deep convolutional networks as shallow gaussian processes”
Adri\‘a Garriga-Alonso, Carl Rasmussen and Laurence Aitchison · 2018
Cited alongside, same era.
“Neural Tangent Kernel: Convergence and Generalization in Neural Networks”
A. Jacot, F. Gabriel and C. Hongler · 2018
Cited alongside, same era.
“Deep Neural Networks as Gaussian Processes”
J. Lee et al · 2018
Cited alongside, same era.
“Introduction to Random Matrices”
Giacomo Livan, Marcel Novaes and Pierpaolo Vivo · 2018
Cited alongside, same era.
“Scaling description of generalization with number of parameters in deep learning”
Mario Geiger et al · 2020
Later among the works it cites.
“Generalisation error in learning with random features and the hidden manifold model”
Federica Gerace et al · 2020
Later among the works it cites.
“Asymptotic errors for convex penalized linear regression beyond Gaussian matrices”
Cedric Gerbelot, Alia Abbara and Florent Krzakala · 2020
Later among the works it cites.
“When Do Neural Networks Outperform Kernel Methods?”
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz and Andrea Montanari · 2020
Later among the works it cites.
“Kernel Alignment Risk Estimator: Risk Prediction from Training Data”
Arthur Jacot et al · 2020
Later among the works it cites.
“Finite versus infinite neural networks: an empirical study”
J. Lee et al · 2020
Later among the works it cites.
“A random matrix analysis of random Fourier features: beyond the Gaussian kernel, a precise phase transition, and the corresponding double descent”
Z. Liao, R. Couillet and Michael. Mahoney · 2020
Later among the works it cites.
“Kernel regression in high dimension: Refined analysis beyond double descent”
F. Liu, Z. Liao and J.A.K. Suykens · 2020
Later among the works it cites.
“Is SGD a Bayesian sampler? Well, almost”
C. Mingard, G. Valle-P\’erez, J. Skalse and A.. Louis · 2020
Later among the works it cites.
“Bayesian Deep Learning and a Probabilistic Perspective of Generalization”
Andrew Wilson and Pavel Izmailov · 2020
Later among the works it cites.
Mikhail Belkin · 2021
Closest in time.
“Spectral bias and task-model alignment explain generalization in kernel regression and infinitely wide neural networks”
Abdulkadir Canatar, Blake Bordelon and Cengiz Pehlevan · 2021
Closest in time.
“Conditioning of Random Feature Matrices: Double Descent and Generalization Error”
Zhijun Chen and Hayden Schaeffer · 2021
Closest in time.
“Towards an Understanding of Benign Overfitting in Neural Networks”
Zhu Li, Zhi-Hua Zhou and Arthur Gretton · 2021
Closest in time.
“Deep double descent: where bigger models and more data hurt”
P. Nakkiran et al · 2021
Closest in time.
D. Bosch, A. Panahi, A. Özcelikkale and D. Dubhash · 2022
Closest in time.
“Generalization error rates in kernel regression: the crossover from the noiseless to noisy regime”
Hugo Cui, Bruno Loureiro, Florent Krzakala and Lenka Zdeborov\’a · 2022
Closest in time.
“Benign, Tempered, or Catastrophic: A Taxonomy of Overfitting”
N. Mallinar et al · 2022
Closest in time.
James. Simon, Madeline Dickens, Dhruva Karkada and Michael. DeWeese · 2022
Closest in time.
Yue. Lu and Horng-Tzer Yau · 2023
Closest in time.