Fetching the paper…
Reading the bibliography…
The relationship between the number of training data points, the number of parameters, and the generalization capabilities of models has been widely studied.
“DISTRIBUTION OF EIGENVALUES FOR SOME SETS OF RANDOM MATRICES”
Vladimir Marcenko and Leonid Pastur · 1967
Earlier work this paper cites.
“Generalized Inversion of Modified Matrices”
Carl. Meyer Jr · 1973
Earlier work this paper cites.
“Training with Noise is Equivalent to Tikhonov Regularization”
Chris. Bishop · 1995
Earlier work this paper cites.
“Statistical Mechanics of Generalization”
Manfred Opper and Wolfgang Kinzel · 1996
Earlier work this paper cites.
“Rate of Convergence to the Semi-Circular Law”
Friedrich Götze and Alexander Tikhomirov · 2003
Earlier work this paper cites.
“Convergence Rates of Spectral Distributions of Large Sample Covariance Matrices”
Z. Bai, Baiqi. Miao and Jian-Feng. Yao · 2003
Earlier work this paper cites.
“Rate of Convergence in Probability to the Marchenko-Pastur Law”
Friedrich Götze and Alexander Tikhomirov · 2004
Earlier work this paper cites.
“The Rate of Convergence for Spectra of GUE and LUE Matrix Ensembles”
Friedrich Götze and Alexander Tikhomirov · 2005
Earlier work this paper cites.
“Eigenvectors of some large sample covariance matrix ensembles”
Olivier Ledoit and Sandrine Péché · 2011
Earlier work this paper cites.
“Understanding the Bias-Variance Tradeoff”, 2012
Scott Fortmann-Roe · 2012
Earlier work this paper cites.
“High-dimensional asymptotics of prediction: Ridge regression and classification”
Edgar Dobriban and Stefan Wager · 2018
Earlier work this paper cites.
“A Mean Field View of the Landscape of Two-layer Neural Networks”
Song Mei, Andrea Montanari and Phan-Minh Nguyen · 2018
Earlier work this paper cites.
“Reconciling Modern Machine-Learning Practice and the Classical Bias–Variance Trade-off”
Mikhail Belkin, Daniel. Hsu, Siyuan Ma and Soumik Mandal · 2019
Earlier work this paper cites.
“Harmless Interpolation of Noisy Data in Regression”
Vidya Muthukumar, Kailas Vodrahalli and Anant Sahai · 2019
Earlier work this paper cites.
“Limitations of Lazy Training of Two-layers Neural Network”
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz and Andrea Montanari · 2019
Earlier work this paper cites.
“A modern maximum-likelihood theory for high-dimensional logistic regression”
Pragya Sur and Emmanuel Candès · 2019
Earlier work this paper cites.
“On the number of variables to use in principal component regression”
Ji Xu and Daniel Hsu · 2019
Earlier work this paper cites.
“PyTorch: An Imperative Style, High-Performance Deep Learning Library”
Adam Paszke et al · 2019
Earlier work this paper cites.
“High-dimensional Dynamics of Generalization Error in Neural Networks”
Madhu. Advani, Andrew. Saxe and Haim Sompolinsky · 2020
Cited alongside, same era.
“Benign Overfitting in Linear Regression”
Peter Bartlett, Philip. Long, Gábor Lugosi and Alexander Tsigler · 2020
Cited alongside, same era.
“Two Models of Double Descent for Weak Features”
Mikhail Belkin, Daniel. Hsu and Ji Xu · 2020
Cited alongside, same era.
“Exact Expressions for Double Descent and Implicit Regularization Via Surrogate Random Design”
Michal Derezinski, Feynman Liang and Michael Mahoney · 2020
Cited alongside, same era.
“Generalisation Error in Learning with Random Features and the Hidden Manifold Model”
Federica Gerace et al · 2020
Cited alongside, same era.
“Kernel and Rich Regimes in Overparametrized Models”
Blake Woodworth et al · 2020
“Learning Curves of Generic Features Maps for Realistic Datasets with a Teacher-Student Model”
Bruno Loureiro et al · 2021
Later among the works it cites.
“Multiple Descent: Design Your Own Generalization Curve”
Lin Chen, Yifei Min, Mikhail Belkin and Amin Karbasi · 2021
Later among the works it cites.
“Linearized Two-layers Neural Networks in High Dimension”
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz and Andrea Montanari · 2021
Later among the works it cites.
“On the Inherent Regularization Effects of Noise Injection During Training”
Oussama Dhifallah and Yue Lu · 2021
Later among the works it cites.
“Eigenvalue distribution of some nonlinear models of random matrices”
Lucas Benigni and Sandrine Péché · 2021
Later among the works it cites.
“Surprises in High-Dimensional Ridgeless Least Squares Interpolation”
Trevor Hastie, Andrea Montanari, Saharon Rosset and Ryan. Tibshirani · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“When Do Neural Networks Outperform Kernel Methods?”
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz and Andrea Montanari · 2020
Cited alongside, same era.
“Optimal Regularization can Mitigate Double Descent”
Preetum Nakkiran, Prayaag Venkat, Sham. Kakade and Tengyu Ma · 2020
Cited alongside, same era.
“The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of Generalization”
Ben Adlam and Jeffrey Pennington · 2020
Cited alongside, same era.
“Triple Descent and the Two Kinds of Overfitting: Where and Why Do They Appear?”
Stéphane d’Ascoli, Levent Sagun and Giulio Biroli · 2020
Cited alongside, same era.
“Analytic study of double descent in binary classification: The impact of loss”
Ganesh Kini and Christos Thrampoulidis · 2020
Cited alongside, same era.
“The role of regularization in classification of high-dimensional noisy gaussian mixture”
Francesca Mignacco et al · 2020
Cited alongside, same era.
Later among the works it cites.
“Dimension Free Ridge Regression”
Chen Cheng and Andrea Montanari · 2022
Later among the works it cites.
“Generalization Error of Random Feature and Kernel Methods: Hypercontractivity and Kernel Matrix Concentration”
Song Mei, Theodor Misiakiewicz and Andrea Montanari · 2022
Later among the works it cites.
“The Optimal Ridge Penalty for Real-World High-Dimensional Data Can Be Zero or Negative Due to the Implicit Ridge Regularization”
Dmitry Kobak, Jonathan Lomond and Benoit Sanchez · 2022
Later among the works it cites.
“Regularization-Wise Double Descent: Why it Occurs and How to Eliminate it”
Fatih Yilmaz and Reinhard Heckel · 2022
Later among the works it cites.
“Precise learning curves and higher-order scaling limits for dot product kernel regression”
Lechao Xiao and Jeffrey Pennington · 2022
Later among the works it cites.
“A model of double descent for high-dimensional binary linear classification”
Zeyu Deng, Abla Kammoun and Christos Thrampoulidis · 2022
Later among the works it cites.
“Dimensionality Reduction, Regularization, and Generalization in Overparameterized Regressions”
Ningyuan Huang, David. Hogg and Soledad Villar · 2022
Later among the works it cites.
“Training Data Size Induced Double Descent For Denoising Feedforward Neural Networks and the Role of Training Noise”
Rishi Sonthalia and Raj Nadakuditi · 2023
Closest in time.
“A U-turn on Double Descent: Rethinking Parameter Counting in Statistical Learning”
Alicia Curth, Alan Jeffares and Mihaela van Schaar · 2023
Closest in time.
“Generalization Error without Independence: Denoising, Linear Regression, and Transfer Learning”, 2023
Chinmaya Kausik, Kashvi Srivastava and Rishi Sonthalia · 2023
Closest in time.
“Near-Interpolators: Rapid Norm Growth and the Trade-Off between Interpolation and Generalization”
Yutong Wang, Rishi Sonthalia and Wei Hu · 2024
Closest in time.