Fetching the paper…
Reading the bibliography…
Neural networks trained by gradient descent (GD) have exhibited a number of surprising generalization behaviors.
“Does data interpolation contradict statistical optimality?”
Mikhail Belkin, Alexander Rakhlin and Alexandre. Tsybakov · 2019
Earlier work this paper cites.
“Probability: theory and examples”
Rick Durrett · 2019
Earlier work this paper cites.
“Surprises in high-dimensional ridgeless least squares interpolation”
Trevor Hastie, Andrea Montanari, Saharon Rosset and Ryan Tibshirani · 2019
Earlier work this paper cites.
“High-Dimensional Statistics: A Non-Asymptotic Viewpoint”, Cambridge Series in Statistical and Probabilistic Mathematics
Martin. Wainwright · 2019
Earlier work this paper cites.
“Regularization Matters: Generalization and Optimization of Neural Nets v.s. their Induced Kernel”
Colin Wei, Jason. Lee, Qiang Liu and Tengyu Ma · 2019
Earlier work this paper cites.
“Benign overfitting in linear regression”
Peter Bartlett, Philip Long, Gábor Lugosi and Alexander Tsigler · 2020
Earlier work this paper cites.
“Just interpolate: Kernel “ridgeless” regression can generalize”
Tengyuan Liang and Alexander Rakhlin · 2020
Earlier work this paper cites.
“Deep learning: a statistical viewpoint”
Peter. Bartlett, Andrea Montanari and Alexander Rakhlin · 2021
Earlier work this paper cites.
“Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation”
Mikhail Belkin · 2021
Earlier work this paper cites.
“Finite-sample Analysis of Interpolating Linear Classifiers in the Overparameterized Regime”
Niladri. Chatterji and Philip. Long · 2021
Earlier work this paper cites.
“Finite-sample analysis of interpolating linear classifiers in the overparameterized regime”
Niladri. Chatterji and Philip. Long · 2021
Earlier work this paper cites.
Yehuda Dar, Vidya Muthukumar and Richard. Baraniuk · 2021
Earlier work this paper cites.
Ke Wang and Christos Thrampoulidis · 2021
Cited alongside, same era.
“Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational Limit”
Boaz Barak, Benjamin. Edelman, Surbhi Goel, Sham Kakade, Eran Malach and Cyril Zhang · 2022
Cited alongside, same era.
“Benign overfitting in two-layer convolutional neural networks”
Yuan Cao, Zixiang Chen, Mikhail Belkin and Quanquan Gu · 2022
Cited alongside, same era.
“Random feature amplification: Feature learning and generalization in neural networks”
Spencer Frei, Niladri Chatterji and Peter Bartlett · 2022
Cited alongside, same era.
“Benign Overfitting without Linearity: Neural Network Classifiers Trained by Gradient Descent for Noisy Linear Data”
“Benign Overfitting in Linear Classifiers and Leaky ReLU Networks from KKT Conditions for Margin Maximization”
Spencer Frei, Gal Vardi, Peter. Bartlett and Nathan Srebro · 2023
Closest in time.
“Implicit Bias in Leaky ReLU Networks Trained on High-Dimensional Data”
Spencer Frei, Gal Vardi, Peter. Bartlett, Nathan Srebro and Wei Hu · 2023
Closest in time.
Andrey Gromov · 2023
Closest in time.
“From Tempered to Benign Overfitting in ReLU Neural Networks”
Guy Kornowski, Gilad Yehudai and Ohad Shamir · 2023
Closest in time.
“Benign Overfitting for Two-layer ReLU Convolutional Networks”
Yiwen Kou, Zixiang Chen, Yuanzhou Chen and Quanquan Gu · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Spencer Frei, Niladri. Chatterji and Peter. Bartlett · 2022
Cited alongside, same era.
“Towards Understanding Grokking: An Effective Theory of Representation Learning”, 2022
Ziming Liu, Ouail Kitouni, Niklas Nolte, Eric. Michaud, Max Tegmark and Mike Williams · 2022
Cited alongside, same era.
“Benign, tempered, or catastrophic: A taxonomy of overfitting”
Neil Mallinar, James Simon, Amirhesam Abedsoltan, Parthe Pandit, Mikhail Belkin and Preetum Nakkiran · 2022
Cited alongside, same era.
“Grokking: Generalization beyond overfitting on small algorithmic datasets”
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin and Vedant Misra · 2022
Cited alongside, same era.
Vimal Thilak, Etai Littwin, Shuangfei Zhai, Omid Saremi, Roni Paiss and Joshua Susskind · 2022
Cited alongside, same era.
“Grokking phase transitions in learning local rules with gradient descent”
Bojan Žunkovič and Enej Ilievski · 2022
Cited alongside, same era.
“Unifying Grokking and Double Descent”, 2023
Xander Davies, Lauro Langosco and David Krueger · 2023
Cited alongside, same era.
“Omnigrok: Grokking Beyond Algorithmic Data”
Ziming Liu, Eric. Michaud and Max Tegmark · 2023
Closest in time.
“A Tale of Two Circuits: Grokking as Competition of Sparse and Dense Subnetworks”, 2023
William Merrill, Nikolaos Tsilivis and Aman Shukla · 2023
Closest in time.
“Progress measures for grokking via mechanistic interpretability”
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith and Jacob Steinhardt · 2023
Closest in time.
“Feature selection and low test error in shallow low-rotation ReLU networks”
Matus Telgarsky · 2023
Closest in time.
“Explaining grokking through circuit efficiency”
Vikrant Varma, Rohin Shah, Zachary Kenton, János Kramár and Ramana Kumar · 2023
Closest in time.
“Benign overfitting of non-smooth neural networks beyond lazy training”
Xingyu Xu and Yuantao Gu · 2023
Closest in time.