Fetching the paper…
Reading the bibliography…
To theoretically understand the behavior of trained deep neural networks, it is necessary to study the dynamics induced by gradient methods from a random initialization.
Barron, a.e.: Universal approximation bounds for superpositions of a sigmoidal function. ieee trans. on information theory 39, 930-945
Andrew Barron · 1993
Earlier work this paper cites.
BAYESIAN LEARNING FOR NEURAL NETWORKS
Radford M Neal · 1995
Earlier work this paper cites.
Priors for infinite networks
Radford M Neal · 1996
Earlier work this paper cites.
Bounds on rates of variable-basis and neural-network approximation
Vera Kurková and Marcello Sanguineti · 2001
Earlier work this paper cites.
A rigorous framework for the mean field limit of multilayer neural networks
Phan-Minh Nguyen and Huy Tuan Pham · 2001
Earlier work this paper cites.
On the tractability of multivariate integration and approximation by neural networks
Hrushikesh Mhaskar · 2003
Earlier work this paper cites.
Convex neural networks
Yoshua Bengio, Nicolas Roux, Pascal Vincent, Olivier Delalleau, and Patrice Marcotte · 2005
Earlier work this paper cites.
A note on the global convergence of multilayer neural networks in the mean field regime
Huy Tuan Pham and Phan-Minh Nguyen · 2006
Earlier work this paper cites.
Tensor programs ii: Neural tangent kernel for any architecture
Greg Yang · 2006
Earlier work this paper cites.
Tensor programs III: neural matrix laws
Greg Yang · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
The dynamics of message passing on dense graphs, with applications to compressed sensing
Mohsen Bayati and Andrea Montanari · 2011
Earlier work this paper cites.
An iterative construction of solutions of the TAP equations for the Sherrington–Kirkpatrick model
Erwin Bolthausen · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2016
Cited alongside, same era.
Bridging nonlinearities and stochastic regularizers with gaussian error linear units
Dan Hendrycks and Kevin Gimpel · 2016
Cited alongside, same era.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Cited alongside, same era.
Stochastic particle gradient descent for infinite ensembles
Atsushi Nitanda and Taiji Suzuki · 2017
Cited alongside, same era.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lénaïc Chizat and Francis Bach · 2018
Cited alongside, same era.
Order and chaos: Ntk views on dnn normalization, checkerboard and boundary artifacts
Arthur Jacot, Franck Gabriel, François Ged, and Clément Hongler · 2019
Later among the works it cites.
Trainability and accuracy of neural networks: An interacting particle system approach, 2019
Grant M. Rotskoff and Eric Vanden-Eijnden · 2019
Later among the works it cites.
Wide feedforward or recurrent neural networks of any architecture are gaussian processes
Greg Yang · 2019
Later among the works it cites.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lenaic Chizat and Francis Bach · 2020
Later among the works it cites.
Batch normalization provably avoids ranks collapse for randomly initialised deep networks
Hadi Daneshmand, Jonas Kohler, Francis Bach, Thomas Hofmann, and Aurelien Lucchi · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Cited alongside, same era.
Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks
Grant Rotskoff and Eric Vanden-Eijnden · 2018
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2019
Cited alongside, same era.
A mean-field limit for certain deep neural networks, 2019
Dyego Araújo, Roberto I. Oliveira, and Daniel Yukimura · 2019
Cited alongside, same era.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Lénaïc Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Modeling from features: a mean-field framework for over-parameterized deep neural networks, 2020
Cong Fang, Jason D. Lee, Pengkun Yang, and Tong Zhang · 2020
Later among the works it cites.
Scaling description of generalization with number of parameters in deep learning
Mario Geiger, Arthur Jacot, Stefano Spigler, Franck Gabriel, Levent Sagun, Stephane d’Ascoli, Giulio Biroli, Clement Hongler, and Matthieu Wyart · 2020
Later among the works it cites.
Disentangling feature and lazy training in deep neural networks
Mario Geiger, Stefano Spigler, Arthur Jacot, and Matthieu Wyart · 2020
Later among the works it cites.
Disentangling feature and lazy training in deep neural networks
Mario Geiger, Stefano Spigler, Arthur Jacot, and Matthieu Wyart · 2020
Later among the works it cites.
Mean field analysis of neural networks: A law of large numbers
Justin Sirignano and Konstantinos Spiliopoulos · 2020
Later among the works it cites.
E. Weinan and Stephan Wojtowytsch · 2020
Later among the works it cites.
Mean field analysis of deep neural networks
Justin Sirignano and Konstantinos Spiliopoulos · 2021
Closest in time.
Tensor programs iv: Feature learning in infinite-width neural networks
Greg Yang and Edward J. Hu · 2021
Closest in time.