Fetching the paper…
Reading the bibliography…
Early stopping is a simple and widely used method to prevent over-training neural networks.
Generalization and parameter estimation in feedforward nets: Some experiments
Nelson Morgan and Hervé Bourlard · 1989
Earlier work this paper cites.
Local operator theory, random matrices and banach spaces
Kenneth R Davidson and Stanislaw J Szarek · 2001
Earlier work this paper cites.
Covering number bounds of certain regularized linear function classes
Tong Zhang · 2002
Earlier work this paper cites.
Optimal regularization can mitigate double descent
Preetum Nakkiran, Prayaag Venkat, Sham Kakade, and Tengyu Ma · 2003
Earlier work this paper cites.
Boosting with early stopping: Convergence and consistency
Tong Zhang and Bin Yu · 2005
Earlier work this paper cites.
On early stopping in gradient descent learning
Yuan Yao, Lorenzo Rosasco, and Andrea Caponnetto · 2007
Earlier work this paper cites.
Multiple descent: Design your own generalization curve
Lin Chen, Yifei Min, Mikhail Belkin, and Amin Karbasi · 2008
Earlier work this paper cites.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
Matrix analysis
Roger A Horn and Charles R Johnson · 2012
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Reconciling modern machine learning practice and the bias-variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2018
Cited alongside, same era.
High-dimensional asymptotics of prediction: Ridge regression and classification
Edgar Dobriban, Stefan Wager, et al · 2018
Cited alongside, same era.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Simon S Du, Wei Hu, and Jason D Lee · 2018
Cited alongside, same era.
Measuring the intrinsic dimension of objective landscapes
Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski · 2018
The implicit regularization of stochastic gradient flow for least squares
Alnur Ali, Edgar Dobriban, and Ryan Tibshirani · 2020
Later among the works it cites.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2020
Later among the works it cites.
Two models of double descent for weak features
Mikhail Belkin, Daniel Hsu, and Ji Xu · 2020
Later among the works it cites.
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
Guy Blanc, Neha Gupta, Gregory Valiant, and Paul Valiant · 2020
Later among the works it cites.
Wonder: Weighted one-shot distributed ridge regression in high dimensions
Edgar Dobriban and Yue Sheng · 2020
Later among the works it cites.
Efficient and scalable bayesian neural nets with rank-1 factors
Michael Dusenberry, Ghassen Jerfel, Yeming Wen, Yian Ma, Jasper Snoek, Katherine Heller, Balaji Lakshminarayanan, and Dustin Tran · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A continuous-time view of early stopping for least squares regression
Alnur Ali, J Zico Kolter, and Ryan J Tibshirani · 2019
Cited alongside, same era.
Exact expressions for double descent and implicit regularization via surrogate random design
Michał Dereziński, Feynman Liang, and Michael W Mahoney · 2019
Cited alongside, same era.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2019
Cited alongside, same era.
The generalization error of random features regression: Precise asymptotics and double descent curve
Song Mei and Andrea Montanari · 2019
Cited alongside, same era.
Partha P Mitra · 2019
Cited alongside, same era.
Andrea Montanari, Feng Ruan, Youngtak Sohn, and Jun Yan · 2019
Cited alongside, same era.
Implicit regularization for optimal sparse recovery
Tomas Vaskevicius, Varun Kanade, and Patrick Rebeschini · 2019
Cited alongside, same era.
Later among the works it cites.
Shape matters: Understanding the implicit bias of the noise covariance
Jeff Z HaoChen, Colin Wei, Jason D Lee, and Tengyu Ma · 2020
Later among the works it cites.
The optimal ridge penalty for real-world high-dimensional data can be zero or negative due to the implicit ridge regularization
Dmitry Kobak, Jonathan Lomond, and Benoit Sanchez · 2020
Later among the works it cites.
Gradient descent with early stopping is provably robust to label noise for overparameterized neural networks
Mingchen Li, Mahdi Soltanolkotabi, and Samet Oymak · 2020
Later among the works it cites.
Harmless interpolation of noisy data in regression
Vidya Muthukumar, Kailas Vodrahalli, Vignesh Subramanian, and Anant Sahai · 2020
Later among the works it cites.
Implicit regularization in deep learning may not be explainable by norms
Noam Razin and Nadav Cohen · 2020
Later among the works it cites.
The statistical complexity of early-stopped mirror descent
Tomas Vaškevičius, Varun Kanade, and Patrick Rebeschini · 2020
Later among the works it cites.