Fetching the paper…
Reading the bibliography…
This paper studies the infinite-width limit of deep linear neural networks initialized with random parameters.
Representations for partially exchangeable arrays of random variables
David J. Aldous · 1981
Earlier work this paper cites.
Approximation and estimation bounds for artificial neural networks
Andrew R. Barron · 1994
Earlier work this paper cites.
Priors for infinite networks
Radford M Neal · 1996
Earlier work this paper cites.
The dynamics of message passing on dense graphs, with applications to compressed sensing
Mohsen Bayati and Andrea Montanari · 2011
Earlier work this paper cites.
An iterative construction of solutions of the TAP equations for the Sherrington–Kirkpatrick model
Erwin Bolthausen · 2014
Earlier work this paper cites.
SGD learns the conjugate kernel class of the network
Amit Daniely · 2017
Earlier work this paper cites.
Stochastic particle gradient descent for infinite ensembles
Atsushi Nitanda and Taiji Suzuki · 2017
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lénaïc Chizat and Francis Bach · 2018
Earlier work this paper cites.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Simon S. Du, Wei Hu, and Jason D. Lee · 2018
Earlier work this paper cites.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nati Srebro · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Earlier work this paper cites.
Neural networks as interacting particle systems: Asymptotic convexity of the loss landscape and universal scaling of the approximation error
Grant M. Rotskoff and Eric Vanden-Eijnden · 2018
Earlier work this paper cites.
High-dimensional probability: An introduction with applications in data science
Roman Vershynin · 2018
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Lénaïc Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Width provably matters in optimization for deep linear neural networks
Simon Du and Wei Hu · 2019
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S. Du, Xiyu Zhai, Barnabás Póczos, and Aarti Singh · 2019
Cited alongside, same era.
Implicit regularization of discrete gradient dynamics in linear neural networks
Gauthier Gidel, Francis Bach, and Simon Lacoste-Julien · 2019
Towards resolving the implicit bias of gradient descent for matrix factorization: Greedy low-rank learning
Zhiyuan Li, Yuping Luo, and Kaifeng Lyu · 2020
Later among the works it cites.
Mean field analysis of neural networks: A law of large numbers
Justin Sirignano and Konstantinos Spiliopoulos · 2020
Later among the works it cites.
On the convergence of gradient descent training for two-layer ReLU-networks in the mean field regime
Stephan Wojtowytsch · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D. Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Later among the works it cites.
Gradient descent on infinitely wide neural networks: Global convergence and generalization
Francis Bach and Lénaïc Chizat · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gradient descent aligns the layers of deep linear networks
Ziwei Ji and Matus Telgarsky · 2019
Cited alongside, same era.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Cited alongside, same era.
A mathematical theory of semantic development in deep neural networks
Andrew M Saxe, James L. McClelland, and Surya Ganguli · 2019
Cited alongside, same era.
Towards a mathematical understanding of neural network-based machine learning: what we know and what we don’t
Weinen E, Chao Ma, Lei Wu, and Stephan Wojtowytsch · 2020
Cited alongside, same era.
Training linear neural networks: Non-local convergence and complexity results
Armin Eftekhari · 2020
Cited alongside, same era.
Disentangling feature and lazy training in deep neural networks
Mario Geiger, Stefano Spigler, Arthur Jacot, and Matthieu Wyart · 2020
Cited alongside, same era.
Training integrable parameterizations of deep neural networks in the infinite-width limit
Karl Hajjar, Lénaïc Chizat, and Christophe Giraud · 2021
Later among the works it cites.
Arthur Jacot, François Ged, Berfin Şimşek, Clément Hongler, and Franck Gabriel · 2021
Later among the works it cites.
Gradient flows on graphons: existence, convergence, continuity equations
Sewoong Oh, Soumik Pal, Raghav Somani, and Raghav Tripathi · 2021
Later among the works it cites.
Implicit bias of SGD for diagonal linear networks: a provable benefit of stochasticity
Scott Pesme, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2021
Later among the works it cites.
Tensor programs IV: Feature learning in infinite-width neural networks
Greg Yang and Edward J. Hu · 2021
Later among the works it cites.
Learning deep linear neural networks: Riemannian gradient flows and convergence to global minimizers
Bubacarr Bah, Holger Rauhut, Ulrich Terstiege, and Michael Westdickenberg · 2022
Closest in time.
The continuous formulation of shallow neural networks as wasserstein-type gradient flows
Xavier Fernández-Real and Alessio Figalli · 2022
Closest in time.
Non-gaussian tensor programs
Eugene Golikov and Greg Yang · 2022
Closest in time.