Fetching the paper…
Reading the bibliography…
This article gives a new proof that fully connected neural networks with random weights and biases converge to Gaussian processes in the regime where the input dimension, output dimension, and depth are kept fixed, while the hidden layer widths tend to infinity.
The organization of behavior; a neuropsycholocigal theory
Donald Olding Hebb · 1949
Earlier work this paper cites.
The perceptron: a probabilistic model for information storage and organization in the brain
Frank Rosenblatt · 1958
Earlier work this paper cites.
On the distribution of the roots of certain symmetric matrices
Eugene P Wigner · 1958
Earlier work this paper cites.
Products of random matrices
Harry Furstenberg and Harry Kesten · 1960
Earlier work this paper cites.
Noncommuting random products
Harry Furstenberg · 1963
Earlier work this paper cites.
Ergodic theory of differentiable dynamical systems
David Ruelle · 1979
Earlier work this paper cites.
Addition of certain non-commuting random variables
Dan Voiculescu · 1986
Earlier work this paper cites.
Priors for infinite networks
Radford M Neal · 1996
Earlier work this paper cites.
Estimation of moments of sums of independent real random variables
Rafał Latała · 1997
Earlier work this paper cites.
Lectures on the combinatorics of free probability
Alexandru Nica and Roland Speicher · 2006
Earlier work this paper cites.
Universal microscopic correlation functions for products of independent ginibre matrices
Gernot Akemann and Zdzislaw Burda · 2012
Earlier work this paper cites.
Products of random matrices: in Statistical Physics
Andrea Crisanti, Giovanni Paladin, and Angelo Vulpiani · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Universal distribution of lyapunov exponents for products of ginibre matrices
Gernot Akemann, Zdzislaw Burda, and Mario Kieburg · 2014
Earlier work this paper cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Earlier work this paper cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Earlier work this paper cites.
Error bounds for approximations with deep relu networks
Dmitry Yarotsky · 2017
Earlier work this paper cites.
Deep convolutional networks as shallow gaussian processes
Adrià Garriga-Alonso, Carl Edward Rasmussen, and Laurence Aitchison · 2018
Earlier work this paper cites.
Gaussian fluctuations for products of random matrices
Vadim Gorin and Yi Sun · 2018
Earlier work this paper cites.
Which neural net architectures give rise to exploding and vanishing gradients?
Boris Hanin · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Gaussian process behaviour in wide deep neural networks
Alexander G de G Matthews, Mark Rowland, Jiri Hron, Richard E Turner, and Zoubin Ghahramani · 2018
Cited alongside, same era.
Bayesian deep convolutional networks with many channels are gaussian processes
Roman Novak, Lechao Xiao, Jaehoon Lee, Yasaman Bahri, Greg Yang, Jiri Hron, Daniel A Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Cited alongside, same era.
The emergence of spectral universality in deep networks
Jeffrey Pennington, Samuel Schoenholz, and Surya Ganguli · 2018
Cited alongside, same era.
Optimal approximation of continuous functions by very deep relu networks
Dmitry Yarotsky · 2018
Cited alongside, same era.
Greg Yang · 2019
Later among the works it cites.
Tensor programs i: Wide feedforward or recurrent neural networks of any architecture are gaussian processes
Greg Yang · 2019
Later among the works it cites.
The neural tangent kernel in high dimensions: Triple descent and a multi-scale theory of generalization
Ben Adlam and Jeffrey Pennington · 2020
Later among the works it cites.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2020
Later among the works it cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ben Adlam, Jake Levinson, and Jeffrey Pennington · 2019
Cited alongside, same era.
Fluctuations of beta-jacobi product processes
Andrew Ahn · 2019
Cited alongside, same era.
From integrable to chaotic systems: Universal local statistics of lyapunov exponents
Gernot Akemann, Zdzislaw Burda, and Mario Kieburg · 2019
Cited alongside, same era.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Cited alongside, same era.
Universal function approximation by deep neural nets with bounded width and relu activations
Boris Hanin · 2019
Cited alongside, same era.
Finite depth and width corrections to the neural tangent kernel
Boris Hanin and Mihai Nica · 2019
Cited alongside, same era.
Later among the works it cites.
Ronald DeVore, Boris Hanin, and Guergana Petrova · 2020
Later among the works it cites.
Spectra of the conjugate kernel and neural tangent kernel for linear-width neural networks
Zhou Fan and Zhichao Wang · 2020
Later among the works it cites.
Dynamics of deep neural networks and neural tangent hierarchy
Jiaoyang Huang and Horng-Tzer Yau · 2020
Later among the works it cites.
On the linearity of large non-linear models: when and why the tangent kernel is constant
Chaoyue Liu, Libin Zhu, and Mikhail Belkin · 2020
Later among the works it cites.
Non-gaussian processes and neural networks at finite widths
Sho Yaida · 2020
Later among the works it cites.
Tensor programs ii: Neural tangent kernel for any architecture
Greg Yang · 2020
Later among the works it cites.
Tensor programs iii: Neural matrix laws
Greg Yang · 2020
Later among the works it cites.
Nonlinear approximation and (deep) relu networks
Ingrid Daubechies, Ronald DeVore, Simon Foucart, Boris Hanin, and Guergana Petrova · 2021
Closest in time.
Non-asymptotic approximations of neural networks by gaussian processes
Ronen Eldan, Dan Mikulincer, and Tselil Schramm · 2021
Closest in time.
Non-asymptotic results for singular values of gaussian matrix products
Boris Hanin and Grigoris Paouris · 2021
Closest in time.
Precise characterization of the prior predictive distribution of deep relu networks
Lorenzo Noci, Gregor Bachmann, Kevin Roth, Sebastian Nowozin, and Thomas Hofmann · 2021
Closest in time.
The principles of deep learning theory
Daniel Roberts, Sho Yaida, and Boris Hanin · 2021
Closest in time.
Exact priors of finite neural networks
Jacob A Zavatone-Veth and Cengiz Pehlevan · 2021
Closest in time.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Closest in time.