Fetching the paper…
Reading the bibliography…
The ability of overparameterized deep networks to interpolate noisy data, while at the same time showing good generalization performance, has been recently characterized in terms of the double descent curve for the test error.
Neural Networks and the Bias/Variance Dilemma
Stuart Geman, Elie Bienenstock, and René Doursat · 1992
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Hessian eigenmaps: Locally linear embedding techniques for high-dimensional data
David L. Donoho and Carrie Grimes · 2003
Earlier work this paper cites.
Sample complexity of testing the manifold hypothesis
Hariharan Narayanan and Sanjoy Mitter · 2010
Earlier work this paper cites.
Wit3: Web inventory of transcribed and translated talks
Mauro Cettolo, Christian Girardi, and Marcello Federico · 2012
Earlier work this paper cites.
Deep learning of representations: Looking forward
Yoshua Bengio · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2014
Earlier work this paper cites.
Results of the wmt14 metrics shared task
Matouš Macháček and Ondřej Bojar · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Convergent learning: Do different neural networks learn the same representations?
Yixuan Li, Jason Yosinski, Jeff Clune, Hod Lipson, and John Hopcroft · 2015
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Overfitting or perfect fitting? risk bounds for classification and regression rules that interpolate
Mikhail Belkin, Daniel J Hsu, and Partha Mitra · 2018
Earlier work this paper cites.
A modern take on the bias-variance tradeoff in neural networks
Brady Neal, Sarthak Mittal, Aristide Baratin, Vinayak Tantia, Matthew Scicluna, Simon Lacoste-Julien, and Ioannis Mitliagkas · 2018
Earlier work this paper cites.
The role of over-parametrization in generalization of neural networks
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro · 2018
Earlier work this paper cites.
Sensitivity and generalization in neural networks: an empirical study
Roman Novak, Yasaman Bahri, Daniel A Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Earlier work this paper cites.
An empirical study of example forgetting during deep neural network learning
Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J Gordon · 2018
Cited alongside, same era.
Curvature-based comparison of two neural networks
Tao Yu, Huan Long, and John E Hopcroft · 2018
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2018
Cited alongside, same era.
How to initialize your network? robust initialization for weightnorm & resnets
Devansh Arpit, Víctor Campos, and Yoshua Bengio · 2019
Cited alongside, same era.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Cited alongside, same era.
Harmless interpolation of noisy data in regression
Vidya Muthukumar, Kailas Vodrahalli, Vignesh Subramanian, and Anant Sahai · 2020
Later among the works it cites.
The intrinsic dimension of images and its impact on learning
Phil Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum, and Tom Goldstein · 2020
Later among the works it cites.
Overfitting in adversarially robust deep learning
Leslie Rice, Eric Wong, and Zico Kolter · 2020
Later among the works it cites.
A case for new neural network smoothness constraints
Mihaela Rosca, Theophane Weber, Arthur Gretton, and Shakir Mohamed · 2020
Later among the works it cites.
Model fusion via optimal transport
Sidak Pal Singh and Martin Jaggi · 2020
Later among the works it cites.
A diffusion theory for deep learning dynamics: Stochastic gradient descent exponentially favors flat minima
Zeke Xie, Issei Sato, and Masashi Sugiyama · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jamming transition as a paradigm to understand the loss landscape of deep neural networks
Mario Geiger, Stefano Spigler, Stéphane d’Ascoli, Levent Sagun, Marco Baity-Jesi, Giulio Biroli, and Matthieu Wyart · 2019
Cited alongside, same era.
Predicting the generalization gap in deep networks with margin distributions
Yiding Jiang, Dilip Krishnan, Hossein Mobahi, and Samy Bengio · 2019
Cited alongside, same era.
Explaining landscape connectivity of low-cost solutions for multilayer nets
Rohith Kuditipudi, Xiang Wang, Holden Lee, Yi Zhang, Zhiyuan Li, Wei Hu, Rong Ge, and Sanjeev Arora · 2019
Cited alongside, same era.
Implicit rugosity regularization via data augmentation
Daniel LeJeune, Randall Balestriero, Hamid Javadi, and Richard G Baraniuk · 2019
Cited alongside, same era.
Robustness via curvature regularization, and vice versa
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, Jonathan Uesato, and Pascal Frossard · 2019
Cited alongside, same era.
Deep double descent, 2019a
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2019
Cited alongside, same era.
Benign overfitting in linear regression
Peter L. Bartlett, Philip M. Long, Gábor Lugosi, and Alexander Tsigler · 2020
Cited alongside, same era.
Later among the works it cites.
Rethinking bias-variance trade-off for generalization of neural networks
Zitong Yang, Yaodong Yu, Chong You, Jacob Steinhardt, and Yi Ma · 2020
Later among the works it cites.
A universal law of robustness via isoperimetry
Sébastien Bubeck and Mark Sellke · 2021
Later among the works it cites.
Provable benefits of overparameterization in model compression: From double descent to pruning neural networks
Xiangyu Chang, Yingcong Li, Samet Oymak, and Christos Thrampoulidis · 2021
Later among the works it cites.
Robust overfitting may be mitigated by properly learned smoothening
Tianlong Chen, Zhenyu Zhang, Sijia Liu, Shiyu Chang, and Zhangyang Wang · 2021
Later among the works it cites.
What happens after sgd reaches zero loss?–a mathematical framework
Zhiyuan Li, Tianhao Wang, and Sanjeev Arora · 2021
Later among the works it cites.
On linear stability of sgd and input-smoothness of neural networks
Chao Ma and Lexing Ying · 2021
Later among the works it cites.
Geometry of the loss landscape in overparameterized neural networks: Symmetries and invariances
Berfin Simsek, François Ged, Arthur Jacot, Francesco Spadaro, Clement Hongler, Wulfram Gerstner, and Johanni Brea · 2021
Later among the works it cites.
Understanding gradient descent on edge of stability in deep learning
Sanjeev Arora, Zhiyuan Li, and Abhishek Panigrahi · 2022
Closest in time.
Are all linear regions created equal?
Matteo Gamba, Adrian Chmielewski-Anders, Josephine Sullivan, Hossein Azizpour, and Mårten Björkman · 2022
Closest in time.
Implicit regularization or implicit conditioning? exact risk trajectories of sgd in high dimensions
Courtney Paquette, Elliot Paquette, Ben Adlam, and Jeffrey Pennington · 2022
Closest in time.
Can neural nets learn the same model twice? investigating reproducibility and double descent from the decision boundary perspective
Gowthami Somepalli, Liam Fowl, Arpit Bansal, Ping Yeh-Chiang, Yehuda Dar, Richard Baraniuk, Micah Goldblum, and Tom Goldstein · 2022
Closest in time.
Gradient-based optimization is not necessary for generalization in neural networks
P. Chiang, R. Ni, D.Ỹ. Miller, A. Bansal, J. Geiping, M. Goldblum, and T. Goldstein · 2023
Closest in time.