Fetching the paper…
Reading the bibliography…
A fundamental problem in machine learning is understanding the effect of early stopping on the parameters obtained and the generalization capabilities of the model.
“Kernel Alignment Risk Estimator: Risk Prediction from Training Data”
Arthur Jacot et al · 2006
Earlier work this paper cites.
“On Early Stopping in Gradient Descent Learning”
Y. Yao, Lorenzo Rosasco and Andrea Caponnetto · 2007
Earlier work this paper cites.
“Random Features for Large-Scale Kernel Machines”
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
“Early stopping and non-parametric regression: An optimal data-dependent stopping rule”
Garvesh Raskutti, Martin Wainwright and Bin Yu · 2013
Earlier work this paper cites.
“Deep Learning” http://www.deeplearningbook.org
Ian Goodfellow, Yoshua Bengio and Aaron Courville · 2016
Earlier work this paper cites.
“High-dimensional asymptotics of prediction: Ridge regression and classification”
Edgar Dobriban and Stefan Wager · 2018
Earlier work this paper cites.
“Neural tangent kernel: Convergence and generalization in neural networks”
Arthur Jacot, Franck Gabriel and Clément Hongler · 2018
Earlier work this paper cites.
“High-dimensional asymptotics of prediction: Ridge regression and classification”
Edgar Dobriban and Stefan Wager · 2018
Earlier work this paper cites.
“A continuous-time view of early stopping for least squares regression”
Alnur Ali, J Kolter and Ryan Tibshirani · 2019
Earlier work this paper cites.
“SGD: General analysis and improved rates”
Robert Gower et al · 2019
Earlier work this paper cites.
“On Lazy Training in Differentiable Programming”
Lenaic Chizat, Edouard Oyallon and Francis Bach · 2019
Earlier work this paper cites.
“PyTorch: An Imperative Style, High-Performance Deep Learning Library”
Adam Paszke et al · 2019
Earlier work this paper cites.
“High-dimensional dynamics of generalization error in neural networks”
Madhu. Advani, Andrew. Saxe and Haim Sompolinsky · 2020
Earlier work this paper cites.
“Optimal Regularization can Mitigate Double Descent”
Preetum Nakkiran, Prayaag Venkat, Sham. Kakade and Tengyu Ma · 2020
Earlier work this paper cites.
“Implicit Regularization of Random Feature Models”
Arthur Jacot et al · 2020
Cited alongside, same era.
“Online robust regression via sgd on the l1 loss”
Scott Pesme and Nicolas Flammarion · 2020
Cited alongside, same era.
“Disentangling feature and lazy training in deep neural networks”
Mario Geiger, Stefano Spigler, Arthur Jacot and Matthieu Wyart · 2020
Cited alongside, same era.
“Benign overfitting in linear regression”
Peter Bartlett, Philip Long, Gábor Lugosi and Alexander Tsigler · 2020
Cited alongside, same era.
“On the Inherent Regularization Effects of Noise Injection During Training”
Oussama Dhifallah and Yue Lu · 2021
Cited alongside, same era.
“Covariate shift in high-dimensional random feature regression”
Nilesh Tripuraneni, Ben Adlam and Jeffrey Pennington · 2021
“Implicit regularization or implicit conditioning? exact risk trajectories of sgd in high dimensions”
Courtney Paquette, Elliot Paquette, Ben Adlam and Jeffrey Pennington · 2022
Later among the works it cites.
Courtney Paquette, Elliot Paquette, Ben Adlam and Jeffrey Pennington · 2022
Later among the works it cites.
“Dimension free ridge regression”
Chen Cheng and Andrea Montanari · 2022
Later among the works it cites.
“Towards Data-Algorithm Dependent Generalization: a Case Study on Overparameterized Linear Regression”
Jing Xu, Jiaye Teng, Yang Yuan and Andrew Yao · 2023
Later among the works it cites.
“Grokking in Linear Estimators–A Solvable Model that Groks without Understanding”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Benign Overfitting of Constant-Stepsize SGD for Linear Regression”
Difan Zou et al · 2021
Cited alongside, same era.
“Early Stopping for Iterative Regularization with General Loss Functions”
Ting Hu and Yunwen Lei · 2022
Cited alongside, same era.
“On Optimal Early Stopping: Over-informative versus Under-informative Parametrization”
Ruoqi Shen, Liyao Gao and Yi-An Ma · 2022
Cited alongside, same era.
“Surprises in High-Dimensional Ridgeless Least Squares Interpolation”
Trevor Hastie, Andrea Montanari, Saharon Rosset and Ryan. Tibshirani · 2022
Cited alongside, same era.
“The Optimal Ridge Penalty for Real-World High-Dimensional Data Can Be Zero or Negative Due to the Implicit Ridge Regularization”
Dmitry Kobak, Jonathan Lomond and Benoit Sanchez · 2022
Cited alongside, same era.
“Regularization-Wise Double Descent: Why it Occurs and How to Eliminate it”
Fatih Yilmaz and Reinhard Heckel · 2022
Cited alongside, same era.
Noam Levi, Alon Beck and Yohai Bar-Sinai · 2023
Later among the works it cites.
“Training Data Size Induced Double Descent For Denoising Feedforward Neural Networks and the Role of Training Noise”
Rishi Sonthalia and Raj Nadakuditi · 2023
Later among the works it cites.
“Under-Parameterized Double Descent for Ridge Regularized Least Squares Denoising of Data on a Line”
Rishi Sonthalia, Xinyue Li and Bochao Gu · 2023
Later among the works it cites.
“Near-interpolators: Rapid norm growth and the trade-off between interpolation and generalization”
Yutong Wang, Rishi Sonthalia and Wei Hu · 2024
Closest in time.
“Double Descent and Overfitting under Noisy Inputs and Distribution Shift for Linear Denoisers”
Chinmaya Kausik, Kashvi Srivastava and Rishi Sonthalia · 2024
Closest in time.
“High-dimensional asymptotics of denoising autoencoders”
Hugo Cui and Lenka Zdeborová · 2024
Closest in time.
“Early alignment in two-layer networks training is a two-edged sword”
Etienne Boursier and Nicolas Flammarion · 2024
Closest in time.
“Stochastic gradient descent for streaming linear and rectified linear systems with Massart noise”
Halyun Jeong, Deanna Needell and Elizaveta Rebrova · 2024
Closest in time.
“Grokking as the transition from lazy to rich training dynamics”
Tanishq Kumar, Blake Bordelon, Samuel. Gershman and Cengiz Pehlevan · 2024
Closest in time.