Fetching the paper…
Reading the bibliography…
We study the distribution of a fully connected neural network with random Gaussian weights and biases in which the hidden layer widths are proportional to a large constant $n$.
The emergence of spectral universality in deep networks
Jeffrey Pennington, Samuel S. Schoenholz, and Surya Ganguli · 1932
Earlier work this paper cites.
Free states of the canonical anticommutation relations
Robert T. Power and Erling Stormers · 1970
Earlier work this paper cites.
Real and complex analysis
Walter Rudin · 1987
Earlier work this paper cites.
On a formula for the l2 wasserstein metric between measures on euclidean and hilbert spaces
Matthias Gelbrich · 1990
Earlier work this paper cites.
On the rate of convergence in the multivariate CLT
F. Götze · 1991
Earlier work this paper cites.
Bayesian learning for neural networks
Radford Neal · 1996
Earlier work this paper cites.
Elements of functional analysis. Transl. from the French by Silvio Levy
Francis Hirsch and Gilles Lacombe · 1999
Earlier work this paper cites.
Sobolev spaces
Robert A. Adams and John J. F. Fournier · 2003
Earlier work this paper cites.
A Lyapunov-type bound in ℝ d \mathbb{R}^{d}
V. Bentkus · 2004
Earlier work this paper cites.
An introduction to infinite-dimensional analysis
Giuseppe Da Prato · 2006
Earlier work this paper cites.
Beyond gaussian processes: on the distributions of infinite networks
Ricky Der and Daniel Lee · 2006
Earlier work this paper cites.
Random fields and geometry
Robert J. Adler and Jonathan E. Taylor · 2007
Earlier work this paper cites.
Level sets and extrema of random processes and fields
Jean-Marc Azais and Mario Wschebor · 2009
Earlier work this paper cites.
Stein’s method on Wiener chaos
Ivan Nourdin and Giovanni Peccati · 2009
Earlier work this paper cites.
Optimal transport
Cédric Villani · 2009
Earlier work this paper cites.
Wiener chaos: Moments, cumulants and diagrams. A survey with computer implementation
Giovanni Peccati and Murad S. Taqqu · 2011
Earlier work this paper cites.
Functional spaces for the theory of elliptic partial differential equations. Transl. from the French by Reinie Erné
Françoise Demengel and Gilbert Demengel · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Normal approximations with Malliavin calculus: from Stein’s method to universality
Ivan Nourdin and Giovanni Peccati · 2012
Earlier work this paper cites.
Matrix analysis
Rajendra Bhatia · 2013
Earlier work this paper cites.
Lectures on Gaussian approximations with Malliavin calculus
Ivan Nourdin · 2013
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Stein’s method, logarithmic Sobolev and transport inequalities
Michel Ledoux, Ivan Nourdin, and Giovanni Peccati · 2015
Earlier work this paper cites.
The optimal fourth moment theorem
Ivan Nourdin and Giovanni Peccati · 2015
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Asymptotic laws for the spatial distribution and the number of connected components of zero sets of Gaussian random functions
F. Nazarov and M. Sodin · 2016
Earlier work this paper cites.
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
Jeffrey Pennington, Samuel Schoenholz, and Surya Ganguli · 2017
Earlier work this paper cites.
Samuel S Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lénaic Chizat and Francis Bach · 2018
Earlier work this paper cites.
Deep convolutional networks as shallow gaussian processes
Adriá Garriga-Alonso, Carl Edward Rasmussen, and Laurence Aitchison · 2018
Earlier work this paper cites.
Which neural net architectures give rise to exploding and vanishing gradients?
Boris Hanin · 2018
Cited alongside, same era.
How to start training: The effect of initialization and architecture
Boris Hanin and David Rolnick · 2018
Cited alongside, same era.
Neural tangent kernel: convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Cited alongside, same era.
Gaussian process behaviour in wide deep neural networks
Alexander Matthews, Jiri Hron, Mark Rowland, Richard Turner, and Zoubin Ghahramani · 2018
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Highly accurate protein structure prediction with alphafold
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al · 2021
Later among the works it cites.
Statistical mechanics of deep linear neural networks: The backpropagating kernel renormalization
Qianyi Li and Haim Sompolinsky · 2021
Later among the works it cites.
A self consistent theory of gaussian processes captures feature learning effects in finite cnns
Gadi Naveh and Zohar Ringel · 2021
Later among the works it cites.
Precise characterization of the prior predictive distribution of deep relu networks
Lorenzo Noci, Gregor Bachmann, Kevin Roth, Sebastian Nowozin, and Thomas Hofmann · 2021
Later among the works it cites.
Mean field analysis of deep neural networks
Justin Sirignano and Konstantinos Spiliopoulos · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Neural networks as interacting particle systems: asymptotic convexity of the loss landscape and universal scaling of the approximation error
Grant Rotskoff and Eric Vanden-Eijnden · 2018
Cited alongside, same era.
A high-dimensional CLT in W 2 W_{2} distance with near optimal convergence rate
Alex Zhai · 2018
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Lénaic Chizat, Edouard OyallonBach, and Francis Bach · 2019
Cited alongside, same era.
Existence of Stein kernels under a spectral gap, and discrepancy bounds
Thomas A. Courtade, Max Fathi, and Ashwin Pananjady · 2019
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Cited alongside, same era.
Tensor programs iib: architectural universality of neural tangent kernel training dynamics
Greg Yang and Etai Littwin · 2021
Later among the works it cites.
Exact marginal prior distributions of finite bayesian neural networks
Jacob Zavatone-Veth and Cengiz Pehlevan · 2021
Later among the works it cites.
Statistical mechanics of deep learning beyond the infinite-width limit
S Ariosto, R Pacelli, M Pastore, F Ginelli, M Gherardi, and P Rotondo · 2022
Later among the works it cites.
High-dimensional asymptotics of feature learning: How one gradient step improves the representation
Jimmy Ba, Murat A Erdogdu, Taiji Suzuki, Zhichao Wang, Denny Wu, and Greg Yang · 2022
Later among the works it cites.
Quantitative convergence of randomly initialized wide deep neural networks towards gaussian processes
Andrea Basteri · 2022
Later among the works it cites.
Quantitative gaussian approximation of randomly initialized deep neural networks
Andrea Basteri and Dario Trevisan · 2022
Later among the works it cites.
Self-consistent dynamical field theory of kernel evolution in wide neural networks
Blake Bordelon and Cengiz Pehlevan · 2022
Later among the works it cites.
Random fully connected neural networks as perturbatively solvable hierarchies
Boris Hanin · 2022
Later among the works it cites.
Rate of convergence of polynomial networks to gaussian processes
Adam Klukowski · 2022
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and the double descent curve
Song Mei and Andrea Montanari · 2022
Later among the works it cites.
Multivariate normal approximation on the wiener space: new bounds in the convex distance
Ivan Nourdin, Giovanni Peccati, and Xiaochuan Yang · 2022
Later among the works it cites.
The principles of deep learning theory
Daniel A Roberts, Sho Yaida, and Boris Hanin · 2022
Later among the works it cites.
Normal approximation of random gaussian neural networks
Nicola Apollonio, Daniela De Canditiis, Giovanni Franzina, Paola Stolfi, and Giovanni Luca Torrisi · 2023
Closest in time.
Krishnakumar Balasubramanian, Larry Goldstein, Nathan Ross, and Adil Salim · 2023
Closest in time.
Non-asymptotic approximations of gaussian neural networks via second-order poincaré inequalities
Alberto Bordino, Stefano Favaro, and Sandra Fortini · 2023
Closest in time.
Optimal learning of deep random networks of extensive-width
Hugo Cui, Florent Krzakala, and Lenka Zdeborová · 2023
Closest in time.
Gage DeZoort and Boris Hanin · 2023
Closest in time.
Small scale clts for the nodal length of monochromatic waves
Gauthier Dierickx, Ivan Nourdin, Giovanni Peccati, and Maurizia Rossi · 2023
Closest in time.
Random neural networks in the infinite width limit as gaussian processes
Boris Hanin · 2023
Closest in time.
Bayesian interpolation with deep linear networks
Boris Hanin and Alexander Zlokapa · 2023
Closest in time.
Vector-valued statistics of binomial processes: Berry-esseen bounds in the convex distance
Mikolaj Kasprzak and Giovanni Peccati · 2023
Closest in time.
Separation of scales and a thermodynamic description of feature learning in some cnns
Inbar Seroussi, Gadi Naveh, and Zohar Ringel · 2023
Closest in time.
Wide deep neural networks with gaussian weights are very close to gaussian processes
Dario Trevisan · 2023
Closest in time.
A quantitative functional central limit theorem for shallow neural networks
Valentina Cammarota, Domenico Marinucci, Michele Salvi, and Stefano Vigogna · 2024
Closest in time.
Critical feature learning in deep neural networks
Kristen Fischer, Javed Lindner, David Dahmen, Zohar Ringel, Michael Kramer, and Mortiz Helias · 2024
Closest in time.