Understand
We consider an existing conjecture addressing the asymptotic behavior of neural networks in the large width limit.
- The results that follow from this conjecture include tight bounds on the behavior of wide networks during stochastic gradient descent, and a derivation of their finite-width dynamics.
- We prove the conjecture for deep networks with polynomial activation functions, greatly extending the validity of these results.
- Finally, we point out a difference in the asymptotic behavior of networks with analytic (and non-linear) activation functions and those with piecewise-linear activations such as ReLU.
Built on
Priors for Infinite Networks , pages 29–53
Radford M. Neal · 1996
Earlier work this paper cites.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nati Srebro · 2017
Earlier work this paper cites.
Reconciling modern machine learning practice and the bias-variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2018
Earlier work this paper cites.
Neural Tangent Kernel: Convergence and Generalization in Neural Networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Similar
Learning curves for deep neural networks: A gaussian field theory perspective
Omry Cohen, Or Malka, and Zohar Ringel · 2019
Cited alongside, same era.
Asymptotics of wide networks from feynman diagrams
Ethan Dyer and Guy Gur-Ari · 2019
Cited alongside, same era.
Finite depth and width corrections to the neural tangent kernel
Boris Hanin and Mihai Nica · 2019
Cited alongside, same era.
Dynamics of deep neural networks and neural tangent hierarchy, 2019
Jiaoyang Huang and Horng-Tzer Yau · 2019
Cited alongside, same era.
Then
Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient Descent
Jaehoon Lee, Lechao Xiao, Samuel S. Schoenholz, Yasaman Bahri, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Later among the works it cites.
Non-gaussian processes and neural networks at finite widths
Sho Yaida · 2019
Later among the works it cites.
Sergey Zagoruyko and Nikos Komodakis · 2019
Later among the works it cites.
On the optimization dynamics of wide hypernetworks
Etai Littwin, Tomer Galanti, and Lior Wolf · 2020
Closest in time.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…