Fetching the paper…
Reading the bibliography…
Supervised deep learning involves the training of neural networks with a large number $N$ of parameters.
Eigenvalues of covariance matrices: Application to neural-network learning
Yann Le Cun, Ido Kanter, and Sara A Solla · 1991
Earlier work this paper cites.
On-line learning in soft committee machines
David Saad and Sara A Solla · 1995
Earlier work this paper cites.
Bayesian Learning for Neural Networks
Radford M. Neal · 1996
Earlier work this paper cites.
Dynamics of training
Siegfried Bös and Manfred Opper · 1997
Earlier work this paper cites.
The mnist database of handwritten digits, 1998
Yann LeCun, Corinna Cortes, and Christopher JC Burges · 1998
Earlier work this paper cites.
A unified bias-variance decomposition
Pedro Domingos · 2000
Earlier work this paper cites.
Statistical mechanics of learning
Andreas Engel and Christian Van den Broeck · 2001
Earlier work this paper cites.
Kernel methods for deep learning
Youngmin Cho and Lawrence K. Saul · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, et al · 2012
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2014
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Earlier work this paper cites.
Universal spectrum of normal modes in low-temperature glasses
Silvio Franz, Giorgio Parisi, Pierfrancesco Urbani, and Francesco Zamponi · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Daniel Soudry and Yair Carmon · 2016
Earlier work this paper cites.
Stuck in a what? adventures in weight space
Zachary C Lipton · 2016
Earlier work this paper cites.
The simplest model of jamming
Silvio Franz and Giorgio Parisi · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Cited alongside, same era.
Topology and geometry of deep rectified network optimization landscapes
C Daniel Freeman and Joan Bruna · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
Singularity of the hessian in deep learning
Levent Sagun, Léon Bottou, and Yann LeCun · 2017
Cited alongside, same era.
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V. Uğur Güney, Yann Dauphin, and Léon Bottou · 2017
Cited alongside, same era.
Energy landscapes for machine learning
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Later among the works it cites.
Just interpolate: Kernel” ridgeless” regression can generalize
Tengyuan Liang and Alexander Rakhlin · 2018
Later among the works it cites.
A note on lazy training in supervised differentiable programming
Lenaic Chizat and Francis Bach · 2018
Later among the works it cites.
Grant M Rotskoff and Eric Vanden-Eijnden · 2018
Later among the works it cites.
A mean field view of the landscape of two-layers neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andrew J Ballard, Ritankar Das, Stefano Martiniani, Dhagash Mehta, Levent Sagun, Jacob D Stevenson, and David J Wales · 2017
Cited alongside, same era.
Geometry of optimization and implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, Ruslan Salakhutdinov, and Nathan Srebro · 2017
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
Madhu S Advani and Andrew M Saxe · 2017
Cited alongside, same era.
Universality of the sat-unsat (jamming) threshold in non-convex continuous constraint satisfaction problems
Silvio Franz, Giorgio Parisi, Maxime Sevelev, Pierfrancesco Urbani, and Francesco Zamponi · 2017
Cited alongside, same era.
Neural networks with finite intrinsic dimension have no spurious valleys
Luca Venturi, Afonso Bandeira, and Joan Bruna · 2018
Cited alongside, same era.
The loss landscape of overparameterized neural networks
Yaim Cooper · 2018
Cited alongside, same era.
Comparing dynamics: Deep neural networks versus glassy systems
Marco Baity-Jesi, Levent Sagun, Mario Geiger, Stefano Spigler, Gerard Ben Arous, Chiara Cammarota, Yann LeCun, Matthieu Wyart, and Giulio Biroli · 2018
Cited alongside, same era.
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Later among the works it cites.
Mean field analysis of neural networks
Justin Sirignano and Konstantinos Spiliopoulos · 2018
Later among the works it cites.
Deep neural networks as gaussian processes
Jae Hoon Lee, Yasaman Bahri, Roman Novak, Samuel S. Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Later among the works it cites.
Jamming in multilayer supervised learning models
Silvio Franz, Sungmin Hwang, and Pierfrancesco Urbani · 2018
Later among the works it cites.
Gradient Descent Provably Optimizes Over-parameterized Neural Networks
Simon S. Du, Xiyu Zhai, Barnabás Póczos, and Aarti Singh · 2019
Closest in time.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 2019
Closest in time.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Closest in time.
Two models of double descent for weak features
Mikhail Belkin, Daniel Hsu, and Ji Xu · 2019
Closest in time.
The generalization error of random features regression: Precise asymptotics and double descent curve
Song Mei and Andrea Montanari · 2019
Closest in time.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2019
Closest in time.
Finite depth and width corrections to the neural tangent kernel
Boris Hanin and Mihai Nica · 2019
Closest in time.
Asymptotics of wide networks from feynman diagrams
Ethan Dyer and Guy Gur-Ari · 2019
Closest in time.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Closest in time.
Mean field limit of the learning dynamics of multilayer neural networks
Phan-Minh Nguyen · 2019
Closest in time.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel S Schoenholz, Yasaman Bahri, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Closest in time.