Fetching the paper…
Reading the bibliography…
Current deep neural networks are highly overparameterized (up to billions of connection weights) and nonlinear.
Bad global minima exist and sgd can reach them
Shengchao Liu, Dimitris Papailiopoulos, and Dimitris Achlioptas · 1906
Earlier work this paper cites.
Fantastic generalization measures and where to find them, 2019
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 1912
Earlier work this paper cites.
Optimization by simulated annealing
S. Kirkpatrick, C. D. Gelatt, and M. P. Vecchi · 1983
Earlier work this paper cites.
Spin glass theory and beyond: An Introduction to the Replica Method and Its Applications
Marc Mézard, Giorgio Parisi, and Miguel Virasoro · 1987
Earlier work this paper cites.
The space of interactions in neural network models
E Gardner · 1988
Earlier work this paper cites.
Optimal storage properties of neural network models
E Gardner and B Derrida · 1988
Earlier work this paper cites.
Storage capacity of memory networks with binary couplings
Werner Krauth and Marc Mézard · 1989
Earlier work this paper cites.
Three unfinished works on the optimal storage capacity of networks
E Gardner and B Derrida · 1989
Earlier work this paper cites.
First-order transition to perfect generalization in a neural network with binary synapses
Géza Györgyi · 1990
Earlier work this paper cites.
Recipes for metastable states in spin glasses
Silvio Franz and Giorgio Parisi · 1995
Earlier work this paper cites.
Bayesian Learning for Neural Networks
Radford M. Neal · 1996
Earlier work this paper cites.
Statistical mechanics of learning
Andreas Engel and Christian Van den Broeck · 2001
Earlier work this paper cites.
Double trouble in double descent : Bias and variance(s) in the lazy regime, 2020
Stèphane d’Ascoli, Maria Refinetti, Giulio Biroli, and Florent Krzakala · 2003
Earlier work this paper cites.
Understanding belief propagation and its generalizations
Jonathan S Yedidia, William T Freeman, Yair Weiss, et al · 2005
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
Efficient supervised learning in networks with binary synapses
Carlo Baldassi, Alfredo Braunstein, Nicolas Brunel, and Riccardo Zecchina · 2007
Earlier work this paper cites.
Universality laws for high-dimensional learning with random features, 2020
Hong Hu and Yue M. Lu · 2009
Earlier work this paper cites.
Generalization learning in a perceptron with binary synapses
Carlo Baldassi · 2009
Earlier work this paper cites.
Memorizing without overfitting: Bias, variance, and interpolation in over-parameterized models
Jason W Rocks and Pankaj Mehta · 2010
Cited alongside, same era.
Mario Geiger, Leonardo Petrini, and Matthieu Wyart · 2012
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Cited alongside, same era.
Origin of the computational hardness for learning with binary synapses
Haiping Huang and Yoshiyuki Kabashima · 2014
Cited alongside, same era.
Subdominant dense clusters allow for simple learning and high computational performance in neural networks with discrete synapses
The generalization error of random features regression: Precise asymptotics and the double descent curve
Song Mei and Andrea Montanari · 2019
Later among the works it cites.
A jamming transition from under- to over-parametrization affects generalization in deep learning
S Spigler, M Geiger, S d’Ascoli, L Sagun, G Biroli, and M Wyart · 2019
Later among the works it cites.
Clustering of solutions in the symmetric binary perceptron
Carlo Baldassi, Riccardo Della Vecchia, Carlo Lucibello, and Riccardo Zecchina · 2020
Later among the works it cites.
Shaping the learning landscape in neural networks around wide flat minima
Carlo Baldassi, Fabrizio Pittorino, and Riccardo Zecchina · 2020
Later among the works it cites.
Wide flat minima and optimal generalization in classifying high-dimensional gaussian mixtures
Carlo Baldassi, Enrico M Malatesta, Matteo Negri, and Riccardo Zecchina · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Carlo Baldassi, Alessandro Ingrosso, Carlo Lucibello, Luca Saglietti, and Riccardo Zecchina · 2015
Cited alongside, same era.
Many-body localization and thermalization in quantum statistical mechanics
Rahul Nandkishore and David A Huse · 2015
Cited alongside, same era.
Unreasonable effectiveness of learning neural networks: From accessible states and robust ensembles to basic algorithmic schemes
Carlo Baldassi, Christian Borgs, Jennifer T. Chayes, Alessandro Ingrosso, Carlo Lucibello, Luca Saglietti, and Riccardo Zecchina · 2016
Cited alongside, same era.
Binarized neural networks
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio · 2016
Cited alongside, same era.
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou · 2017
Cited alongside, same era.
Efficiency of quantum vs. classical annealing in nonconvex learning problems
Carlo Baldassi and Riccardo Zecchina · 2018
Cited alongside, same era.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S. Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Cited alongside, same era.
Modeling the influence of data structure on learning in neural networks: The hidden manifold model
Sebastian Goldt, Marc Mézard, Florent Krzakala, and Lenka Zdeborová · 2020
Later among the works it cites.
Generalisation error in learning with random features and the hidden manifold model
Federica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mezard, and Lenka Zdeborova · 2020
Later among the works it cites.
Scaling description of generalization with number of parameters in deep learning
Mario Geiger, Arthur Jacot, Stefano Spigler, Franck Gabriel, Levent Sagun, Stéphane d’Ascoli, Giulio Biroli, Clément Hongler, and Matthieu Wyart · 2020
Later among the works it cites.
Unveiling the structure of wide flat minima in neural networks
Carlo Baldassi, Clarissa Lauditi, Enrico M Malatesta, Gabriele Perugini, and Riccardo Zecchina · 2021
Closest in time.
Frozen 1-rsb structure of the symmetric ising perceptron
Will Perkins and Changji Xu · 2021
Closest in time.
Proof of the contiguity conjecture and lognormal limit for the symmetric perceptron
Emmanuel Abbe, Shuangping Li, and Allan Sly · 2021
Closest in time.
Binary perceptron: efficient algorithms can find solutions in a rare well-connected cluster
Emmanuel Abbe, Shuangping Li, and Allan Sly · 2021
Closest in time.
The inverse variance–flatness relation in stochastic gradient descent is critical for finding flat minima
Yu Feng and Yuhai Tu · 2021
Closest in time.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2021
Closest in time.
Entropic gradient descent algorithms and wide flat minima
Fabrizio Pittorino, Carlo Lucibello, Christoph Feinauer, Gabriele Perugini, Carlo Baldassi, Elizaveta Demyanenko, and Riccardo Zecchina · 2021
Closest in time.
Entropic gradient descent algorithms and wide flat minima
Fabrizio Pittorino, Carlo Lucibello, Christoph Feinauer, Gabriele Perugini, Carlo Baldassi, Elizaveta Demyanenko, and Riccardo Zecchina · 2021
Closest in time.
The gaussian equivalence of generative models for learning with shallow neural networks
Sebastian Goldt, Bruno Loureiro, Galen Reeves, Florent Krzakala, Marc Mezard, and Lenka Zdeborova · 2022
Closest in time.
Fabrizio Pittorino, Antonio Ferraro, Gabriele Perugini, Christoph Feinauer, Carlo Baldassi, and Riccardo Zecchina · 2022
Closest in time.