Fetching the paper…
Reading the bibliography…
Two distinct limits for deep learning have been derived as the network width $h\rightarrow \infty$, depending on how the weights of the last layer scale with $h$.
Bayesian Learning for Neural Networks
Radford M. Neal · 1996
Earlier work this paper cites.
Computing with infinite networks
Christopher KI Williams · 1997
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
High-dimensional dynamics of generalization error in neural networks
Madhu S Advani and Andrew M Saxe · 2017
Earlier work this paper cites.
Geometry of optimization and implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, Ruslan Salakhutdinov, and Nathan Srebro · 2017
Earlier work this paper cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms, 2017
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2018
Earlier work this paper cites.
Comparing dynamics: Deep neural networks versus glassy systems
Marco Baity-Jesi, Levent Sagun, Mario Geiger, Stefano Spigler, Gerard Ben Arous, Chiara Cammarota, Yann LeCun, Matthieu Wyart, and Giulio Biroli · 2018
Earlier work this paper cites.
Minnorm training: an algorithm for training overcomplete deep neural networks
Yamini Bansal, Madhu Advani, David D Cox, and Andrew M Saxe · 2018
Earlier work this paper cites.
On the Global Convergence of Gradient Descent for Over-parameterized Models using Optimal Transport
Lénaïc Chizat and Francis Bach · 2018
Earlier work this paper cites.
Gaussian process behaviour in wide deep neural networks
Alexander G. de G. Matthews, Jiri Hron, Mark Rowland, Richard E. Turner, and Zoubin Ghahramani · 2018
Earlier work this paper cites.
The jamming transition as a paradigm to understand the loss landscape of deep neural networks
Mario Geiger, Stefano Spigler, Stéphane d’Ascoli, Levent Sagun, Marco Baity-Jesi, Giulio Biroli, and Matthieu Wyart · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Deep neural networks as gaussian processes
Jae Hoon Lee, Yasaman Bahri, Roman Novak, Samuel S. Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Cited alongside, same era.
A mean field view of the landscape of two-layers neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Cited alongside, same era.
Asymptotics of wide networks from feynman diagrams
Ethan Dyer and Guy Gur-Ari · 2019
Closest in time.
Scaling description of generalization with number of parameters in deep learning
Mario Geiger, Arthur Jacot, Stefano Spigler, Franck Gabriel, Levent Sagun, Stéphane d’Ascoli, Giulio Biroli, Clément Hongler, and Matthieu Wyart · 2019
Closest in time.
The asymptotic spectrum of the hessian of dnn throughout training
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2019
Closest in time.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel S Schoenholz, Yasaman Bahri, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Closest in time.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Grant M Rotskoff and Eric Vanden-Eijnden · 2018
Cited alongside, same era.
Mean field analysis of neural networks
Justin Sirignano and Konstantinos Spiliopoulos · 2018
Cited alongside, same era.
A jamming transition from under-to over-parametrization affects loss landscape and generalization
Stefano Spigler, Mario Geiger, Stéphane d’Ascoli, Levent Sagun, Giulio Biroli, and Matthieu Wyart · 2018
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 2019
Cited alongside, same era.
A Note on Lazy Training in Supervised Differentiable Programming
Lenaic Chizat and Francis Bach · 2019
Cited alongside, same era.
On Lazy Training in Differentiable Programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Cited alongside, same era.
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Closest in time.
A modern take on the bias-variance tradeoff in neural networks
Brady Neal, Sarthak Mittal, Aristide Baratin, Vinayak Tantia, Matthew Scicluna, Simon Lacoste-Julien, and Ioannis Mitliagkas · 2019
Closest in time.
Mean field limit of the learning dynamics of multilayer neural networks
Phan-Minh Nguyen · 2019
Closest in time.
Bayesian deep convolutional networks with many channels are gaussian processes
Roman Novak, Lechao Xiao, Yasaman Bahri, Jaehoon Lee, Greg Yang, Daniel A. Abolafia, Jeffrey Pennington, and Jascha Sohl-dickstein · 2019
Closest in time.
The effect of network width on stochastic gradient descent and generalization: an empirical study
Daniel S Park, Jascha Sohl-Dickstein, Quoc V Le, and Samuel L Smith · 2019
Closest in time.
Greg Yang · 2019
Closest in time.
Geometric compression of invariant manifolds in neural nets, 2020
Jonas Paccolat, Leonardo Petrini, Mario Geiger, Kevin Tyloo, and Matthieu Wyart · 2020
Closest in time.