Fetching the paper…
Reading the bibliography…
We argue that in fully-connected networks a phase transition delimits the over- and under-parametrized regimes where fitting can or cannot be achieved.
A simple weight decay can improve generalization
Anders Krogh and John A Hertz · 1992
Earlier work this paper cites.
Convolutional networks for images, speech, and time series
Yann LeCun, Yoshua Bengio, et al · 1995
Earlier work this paper cites.
On-line learning in soft committee machines
David Saad and Sara A Solla · 1995
Earlier work this paper cites.
Early stopping-but when?
Lutz Prechelt · 1998
Earlier work this paper cites.
Stress propagation through frictionless granular material
Alexei V. Tkachenko and Thomas A. Witten · 1999
Earlier work this paper cites.
Overfitting in neural nets: Backpropagation, conjugate gradient, and early stopping
Rich Caruana, Steve Lawrence, and C Lee Giles · 2001
Earlier work this paper cites.
Statistical mechanics of learning
Andreas Engel and Christian Van den Broeck · 2001
Earlier work this paper cites.
On the rigidity of amorphous solids
M. Wyart · 2005
Earlier work this paper cites.
Effects of compression on the vibrational modes of marginally jammed solids
Matthieu Wyart, Leonardo E Silbert, Sidney R Nagel, and Thomas A Witten · 2005
Earlier work this paper cites.
The jamming scenario - an introduction and outlook
Andrea J Liu, Sidney R Nagel, W Saarloos, and Matthieu Wyart · 2010
Earlier work this paper cites.
Theoretical perspective on the glass transition and amorphous materials
Ludovic Berthier and Giulio Biroli · 2011
Earlier work this paper cites.
Phonon gap and localization lengths in floppy materials
Gustavo Düring, Edan Lerner, and Matthieu Wyart · 2013
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Cited alongside, same era.
Universal spectrum of normal modes in low-temperature glasses
Silvio Franz, Giorgio Parisi, Pierfrancesco Urbani, and Francesco Zamponi · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Geometry of optimization and implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, Ruslan Salakhutdinov, and Nathan Srebro · 2017
Later among the works it cites.
High-dimensional dynamics of generalization error in neural networks
Madhu S Advani and Andrew M Saxe · 2017
Later among the works it cites.
Neural networks with finite intrinsic dimension have no spurious valleys
Luca Venturi, Afonso Bandeira, and Joan Bruna · 2018
Closest in time.
The loss landscape of overparameterized neural networks
Yaim Cooper · 2018
Closest in time.
Comparing dynamics: Deep neural networks versus glassy systems
Marco Baity-Jesi, Levent Sagun, Mario Geiger, Stefano Spigler, Gerard Ben Arous, Chiara Cammarota, Yann LeCun, Matthieu Wyart, and Giulio Biroli · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daniel Soudry and Yair Carmon · 2016
Cited alongside, same era.
Stuck in a what? adventures in weight space
Zachary C Lipton · 2016
Cited alongside, same era.
Samuel S Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein · 2016
Cited alongside, same era.
The simplest model of jamming
Silvio Franz and Giorgio Parisi · 2016
Cited alongside, same era.
Topology and geometry of deep rectified network optimization landscapes
C Daniel Freeman and Joan Bruna · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
Singularity of the hessian in deep learning
Levent Sagun, Léon Bottou, and Yann LeCun · 2017
Cited alongside, same era.
Closest in time.
The jamming transition as a paradigm to understand the loss landscape of deep neural networks
Mario Geiger, Stefano Spigler, Stéphane d’Ascoli, Levent Sagun, Marco Baity-Jesi, Giulio Biroli, and Matthieu Wyart · 2018
Closest in time.
Towards understanding the role of over-parametrization in generalization of neural networks
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro · 2018
Closest in time.
Minnorm training: an algorithm for training over-parameterized deep neural networks
Yamini Bansal, Madhu Advani, David D Cox, and Andrew M Saxe · 2018
Closest in time.
Reconciling modern machine learning and the bias-variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2018
Closest in time.
The dynamics of learning: A random matrix approach
Zhenyu Liao and Romain Couillet · 2018
Closest in time.
Jamming in multilayer supervised learning models
P. Urbani S. Franz, S. Hwang · 2018
Closest in time.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2019
Closest in time.
Two models of double descent for weak features
Mikhail Belkin, Daniel Hsu, and Ji Xu · 2019
Closest in time.
Scaling description of generalization with number of parameters in deep learning
Mario Geiger, Arthur Jacot, Stefano Spigler, Franck Gabriel, Levent Sagun, Stéphane d’Ascoli, Giulio Biroli, Clément Hongler, and Matthieu Wyart · 2019
Closest in time.