Fetching the paper…
Reading the bibliography…
The dynamics of DNNs during gradient descent is described by the so-called Neural Tangent Kernel (NTK).
Scaling description of generalization with number of parameters in deep learning
Mario Geiger, Arthur Jacot, Stefano Spigler, Franck Gabriel, Levent Sagun, Stéphane d’Ascoli, Giulio Biroli, Clément Hongler, and Matthieu Wyart · 1901
Earlier work this paper cites.
Vardan Papyan · 1901
Earlier work this paper cites.
Freeze and chaos for dnns: an NTK view of batch normalization, checkerboard and boundary effects
Arthur Jacot, Franck Gabriel, and Clément Hongler · 1907
Earlier work this paper cites.
Bayesian Learning for Neural Networks
Radford M. Neal · 1996
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Information geometry of neural networks
Daniel Wagenaar · 1998
Earlier work this paper cites.
Kernel Methods for Deep Learning
Youngmin Cho and Lawrence K. Saul · 2009
Earlier work this paper cites.
Revisiting Natural Gradient for Deep Networks
Razvan Pascanu and Yoshua Bengio · 2013
Earlier work this paper cites.
Identifying and Attacking the Saddle Point Problem in High-dimensional Non-convex Optimization
Yann N. Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Earlier work this paper cites.
On the saddle point problem for non-convex optimization
Razvan Pascanu, Yann N Dauphin, Surya Ganguli, and Yoshua Bengio · 2014
Earlier work this paper cites.
The Loss Surfaces of Multilayer Networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Earlier work this paper cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2016
Cited alongside, same era.
Exponential expressivity in deep neural networks through transient chaos
Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli · 2016
Cited alongside, same era.
Singularity of the hessian in deep learning
Levent Sagun, Léon Bottou, and Yann LeCun · 2016
Cited alongside, same era.
Geometry of Neural Network Loss Surfaces via Random Matrix Theory
Jeffrey Pennington and Yasaman Bahri · 2017
Cited alongside, same era.
Empirical Analysis of the Hessian of Over-Parametrized Neural Networks
Levent Sagun, Utku Evci, V. Ugur Güney, Yann Dauphin, and Léon Bottou · 2017
Deep Neural Networks as Gaussian Processes
Jae Hoon Lee, Yasaman Bahri, Roman Novak, Samuel S. Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Later among the works it cites.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Later among the works it cites.
The Spectrum of the Fisher Information Matrix of a Single-Hidden-Layer Neural Network
Jeffrey Pennington and Pratik Worah · 2018
Later among the works it cites.
Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks
Grant Rotskoff and Eric Vanden-Eijnden · 2018
Later among the works it cites.
How does batch normalization help optimization?
Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry · 2018
Later among the works it cites.
Towards Understanding Generalization of Deep Learning: Perspective of Loss Landscapes
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A Convergence Theory for Deep Learning via Over-Parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2018
Cited alongside, same era.
Comparing Dynamics: Deep Neural Networks versus Glassy Systems
Marco Baity-Jesi, Levent Sagun, Mario Geiger, Stefano Spigler, Gerard Ben Arous, Chiara Cammarota, Yann LeCun, Matthieu Wyart, and Giulio Biroli · 2018
Cited alongside, same era.
The jamming transition as a paradigm to understand the loss landscape of deep neural networks
Mario Geiger, Stefano Spigler, Stéphane d’Ascoli, Levent Sagun, Marco Baity-Jesi, Giulio Biroli, and Matthieu Wyart · 2018
Cited alongside, same era.
Gradient descent happens in a tiny subspace
Guy Gur-Ari, Daniel A. Roberts, and Ethan Dyer · 2018
Cited alongside, same era.
Neural Tangent Kernel: Convergence and Generalization in Neural Networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Universal Statistics of Fisher Information in Deep Neural Networks: Mean Field Approach
Ryo Karakida, Shotaro Akaho, and Shun-Ichi Amari · 2018
Cited alongside, same era.
On the Global Convergence of Gradient Descent for Over-parameterized Models using Optimal Transport
Lénaïc Chizat and Francis Bach
Cited in the paper.
Lei Wu, Zhanxing Zhu, and Weinan E · 2018
Later among the works it cites.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 2019
Closest in time.
Gradient Descent Provably Optimizes Over-parameterized Neural Networks
Simon S. Du, Xiyu Zhai, Barnabás Póczos, and Aarti Singh · 2019
Closest in time.
An investigation into neural net optimization via hessian eigenvalue density
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao · 2019
Closest in time.
Dynamics of deep neural networks and neural tangent hierarchy
Jiaoyang Huang and Horng-Tzer Yau · 2019
Closest in time.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Closest in time.