Fetching the paper…
Reading the bibliography…
One of the central puzzles in modern machine learning is the ability of heavily overparametrized models to generalize well.
Spin glass theory and beyond: An Introduction to the Replica Method and Its Applications
Marc Mézard, Giorgio Parisi, and Miguel Virasoro · 1987
Earlier work this paper cites.
Statistical mechanics of generalization
Manfred Opper and Wolfgang Kinzel · 1996
Earlier work this paper cites.
Statistical mechanics of learning
Andreas Engel and Christian Van den Broeck · 2001
Earlier work this paper cites.
Landscape analysis of constraint satisfaction problems
Florent Krzakala and Jorge Kurchan · 2007
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
The spectrum of kernel random matrices
Noureddine El Karoui et al · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, et al · 2012
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
High-dimensional asymptotics of prediction: Ridge regression and classification
Edgar Dobriban and Stefan Wager · 2015
Earlier work this paper cites.
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
The simplest model of jamming
Silvio Franz and Giorgio Parisi · 2016
Earlier work this paper cites.
Towards understanding the role of over-parametrization in generalization of neural networks
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro · 2018
Earlier work this paper cites.
Reconciling modern machine learning and the bias-variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2018
Earlier work this paper cites.
The committee machine: Computational to statistical gaps in learning a two-layers neural network
Benjamin Aubin, Antoine Maillard, Jean Barbier, Florent Krzakala, Nicolas Macris, and Lenka Zdeborová · 2018
Earlier work this paper cites.
Entropy and mutual information in models of deep neural networks
Marylou Gabrié, Andre Manoel, Clément Luneau, Jean Barbier, Nicolas Macris, Florent Krzakala, and Lenka Zdeborová · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2019
Cited alongside, same era.
A jamming transition from under-to over-parametrization affects generalization in deep learning
S Spigler, M Geiger, S d’Ascoli, L Sagun, G Biroli, and M Wyart · 2019
Cited alongside, same era.
Jamming transition as a paradigm to understand the loss landscape of deep neural networks
Mario Geiger, Stefano Spigler, Stéphane d’Ascoli, Levent Sagun, Marco Baity-Jesi, Giulio Biroli, and Matthieu Wyart · 2019
Cited alongside, same era.
The generalization error of random features regression: Precise asymptotics and double descent curve
Song Mei and Andrea Montanari · 2019
Cited alongside, same era.
Optimal errors and phase transitions in high-dimensional generalized linear models
Jean Barbier, Florent Krzakala, Nicolas Macris, Léo Miolane, and Lenka Zdeborová · 2019
The gaussian equivalence of generative models for learning with two-layer neural networks
Sebastian Goldt, Galen Reeves, Marc Mézard, Florent Krzakala, and Lenka Zdeborová · 2020
Later among the works it cites.
Universality laws for high-dimensional learning with random features
Hong Hu and Yue M Lu · 2020
Later among the works it cites.
Optimal regularization can mitigate double descent
Preetum Nakkiran, Prayaag Venkat, Sham Kakade, and Tengyu Ma · 2020
Later among the works it cites.
On the optimal weighted l2 regularization in overparameterized linear regression
Denny Wu and Ji Xu · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Optimal ridge penalty for real-world high-dimensional data can be zero or negative due to the implicit ridge regularization
Dmitry Kobak, Jonathan Lomond, and Benoit Sanchez · 2019
Cited alongside, same era.
Andrea Montanari, Feng Ruan, Youngtak Sohn, and Jun Yan · 2019
Cited alongside, same era.
A model of double descent for high-dimensional binary linear classification
Zeyu Deng, Abla Kammoun, and Christos Thrampoulidis · 2019
Cited alongside, same era.
Asymptotic learning curves of kernel methods: empirical data vs teacher-student paradigm
Stefano Spigler, Mario Geiger, and Matthieu Wyart · 2019
Cited alongside, same era.
Intrinsic dimension of data representations in deep neural networks
Alessio Ansuini, Alessandro Laio, Jakob H Macke, and Davide Zoccolan · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Lénaïc Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Harnessing the power of infinitely wide deep nets on small-data tasks
Sanjeev Arora, Simon S Du, Zhiyuan Li, Ruslan Salakhutdinov, Ruosong Wang, and Dingli Yu · 2019
Cited alongside, same era.
Lin Chen, Yifei Min, Mikhail Belkin, and Amin Karbasi · 2020
Later among the works it cites.
When do neural networks outperform kernel methods?
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2020
Later among the works it cites.
Zhenyu Liao, Romain Couillet, and Michael W Mahoney · 2020
Later among the works it cites.
Statistical mechanics of generalization in kernel regression
Abdulkadir Canatar, Blake Bordelon, and Cengiz Pehlevan · 2020
Later among the works it cites.
Finite-sample analysis of interpolating linear classifiers in the overparameterized regime
Niladri S Chatterji and Philip M Long · 2020
Later among the works it cites.
A precise high-dimensional asymptotic theory for boosting and min-l1-norm interpolated classifiers
Tengyuan Liang and Pragya Sur · 2020
Later among the works it cites.
The role of regularization in classification of high-dimensional noisy gaussian mixture
Francesca Mignacco, Florent Krzakala, Yue Lu, Pierfrancesco Urbani, and Lenka Zdeborova · 2020
Later among the works it cites.
Analytic study of double descent in binary classification: The impact of loss
Ganesh Ramachandra Kini and Christos Thrampoulidis · 2020
Later among the works it cites.
Classification vs regression in overparameterized regimes: Does the loss function matter?
Vidya Muthukumar, Adhyyan Narang, Vignesh Subramanian, Mikhail Belkin, Daniel Hsu, and Anant Sahai · 2020
Later among the works it cites.
Implicit regularization in deep learning: A view from function space
Aristide Baratin, Thomas George, César Laurent, R Devon Hjelm, Guillaume Lajoie, Pascal Vincent, and Simon Lacoste-Julien · 2020
Later among the works it cites.
Evaluation of neural architectures trained with square loss vs cross-entropy in classification tasks
Like Hui and Mikhail Belkin · 2020
Later among the works it cites.
Exploring the role of loss functions in multiclass classification
Ahmet Demirkaya, Jiasi Chen, and Samet Oymak · 2020
Later among the works it cites.
Triple descent and the two kinds of overfitting: where and why do they appear?
Stéphane d'Ascoli, Levent Sagun, and Giulio Biroli · 2020
Later among the works it cites.
The intrinsic dimension of images and its impact on learning
Phillip Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum, and Tom Goldstein · 2021
Closest in time.
Capturing the learning curves of generic features maps for realistic data sets with a teacher-student model
Bruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt, Florent Krzakala, Marc Mézard, Lenka Zdeborová, École Fédérale, and Polytechnique De Lausanne · 2021
Closest in time.