Fetching the paper…
Reading the bibliography…
We show that taking the width and depth to infinity in a deep neural network with skip connections, when branches are scaled by $1/\sqrt{depth}$ (the only nontrivial scaling), result in the same covariance structure no matter how that limit is taken.
“Classical Fifth-, Sixth-, Seventh-, and Eighth-Order Runge-Kutta Formulas with Stepsize Control”
E. Fehlberg · 1968
Earlier work this paper cites.
“Theory of Financial Decision Making”
Jonathan. Ingersoll · 1987
Earlier work this paper cites.
“Numerical Solution of Stochastic Differential Equations”
Peter Kloeden and Eckhard Platen · 1995
Earlier work this paper cites.
“Bayesian Learning for Neural Networks”
R.M. Neal · 1995
Earlier work this paper cites.
“Stochastic Differential Equations”
Bernt Øksendal · 2003
Earlier work this paper cites.
“Tensor Programs II: Neural Tangent Kernel for Any Architecture” arXiv: 2006.14548
Greg Yang · 2006
Earlier work this paper cites.
“Nonlinear SDEs driven by Lévy processes and related PDEs”
Benjamin Jourdain, Sylvie Meleard and Wojbor Woyczynski · 2007
Earlier work this paper cites.
“Stochastic differential equations and applications”
Mao Xuerong · 2008
Earlier work this paper cites.
“Exponential expressivity in deep neural networks through transient chaos”
B. Poole et al · 2016
Earlier work this paper cites.
“Deep Information Propagation”
S.S. Schoenholz, J. Gilmer, S. Ganguli and J. Sohl-Dickstein · 2017
Earlier work this paper cites.
“Mean field residual networks: On the edge of chaos”
G. Yang and S. Schoenholz · 2017
Earlier work this paper cites.
“Deep Neural Networks as Gaussian Processes”
J. Lee et al · 2018
Earlier work this paper cites.
“Gaussian Process Behaviour in Wide Deep Neural Networks”
A.G. Matthews et al · 2018
Earlier work this paper cites.
“CALCUL STOCHASTIQUE ET FINANCE”, 2018
Peter Tankov and Nizar Touzi · 2018
Earlier work this paper cites.
“Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks”
Sanjeev Arora et al · 2019
Earlier work this paper cites.
“Universal Function Approximation by Deep Neural Nets with Bounded Width and ReLU Activations”
Boris Hanin · 2019
Earlier work this paper cites.
“Products of Many Large Random Matrices and Gradients in Deep Neural Networks”
Boris Hanin and Mihai Nica · 2019
Earlier work this paper cites.
“On the Impact of the Activation Function on Deep Neural Networks Training”
S. Hayou, A. Doucet and J. Rousseau · 2019
Earlier work this paper cites.
“Training dynamics of deep networks using stochastic gradient descent via neural tangent kernel”
Soufiane Hayou, Arnaud Doucet and Judith Rousseau · 2019
Cited alongside, same era.
“Understanding Priors in Bayesian Neural Networks at the Unit Level”
Mariia Vladimirova, Jakob Verbeek, Pablo Mesejo and Julyan Arbel · 2019
Cited alongside, same era.
G. Yang · 2019
Cited alongside, same era.
G. Yang · 2019
Cited alongside, same era.
“A fine-grained spectral perspective on neural networks”
Greg Yang and Hadi Salman · 2019
“Tensor Programs IV: Feature Learning in Infinite-Width Neural Networks”
G. Yang and E.J. Hu · 2021
Later among the works it cites.
Greg Yang and Etai Littwin · 2021
Later among the works it cites.
“Exact marginal prior distributions of finite Bayesian neural networks”
Jacob Zavatone-Veth and Cengiz Pehlevan · 2021
Later among the works it cites.
“Correlation Functions in Random Fully Connected Neural Networks at Finite Width”
Boris Hanin · 2022
Later among the works it cites.
“On the infinite-depth limit of finite-width neural networks”
Soufiane Hayou · 2022
Later among the works it cites.
“The Curse of Depth in Kernel Regime”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Finite Depth and Width Corrections to the Neural Tangent Kernel”
Boris Hanin and Mihai Nica · 2020
Cited alongside, same era.
“Mean-field Behaviour of Neural Tangent Kernel for Deep Neural Networks”
S. Hayou, A. Doucet and J. Rousseau · 2020
Cited alongside, same era.
“Bayesian Deep Ensembles via the Neural Tangent Kernel”
Bobby He, Balaji Lakshminarayanan and Yee Teh · 2020
Cited alongside, same era.
“Infinite attention: NNGP and NTK for deep attention networks”
Jiri Hron, Yasaman Bahri, Jascha Sohl-Dickstein and Roman Novak · 2020
Cited alongside, same era.
“Infinitely deep neural networks as diffusion processes”
Stefano Peluchetti and Stefano Favaro · 2020
Cited alongside, same era.
“Disentangling Trainability and Generalization in Deep Neural Networks”
Lechao Xiao, Jeffrey Pennington and Samuel Schoenholz · 2020
Cited alongside, same era.
“Tensor Programs III: Neural Matrix Laws”
G. Yang · 2020
Cited alongside, same era.
Soufiane Hayou, Arnaud Doucet and Judith Rousseau · 2022
Later among the works it cites.
“Theory of Deep Learning: Neural Tangent Kernel and Beyond”
Arthur Jacot · 2022
Later among the works it cites.
“Freeze and Chaos: NTK views on DNN Normalization, Checkerboard and Boundary Artifacts”
Arthur Jacot, Franck Gabriel, Francois Ged and Clement Hongler · 2022
Later among the works it cites.
“The Neural Covariance SDE: Shaped Infinite Depth-and-Width Networks at Initialization”
Mufan Li, Mihai Nica and Daniel. Roy · 2022
Later among the works it cites.
“Connecting Optimization and Generalization via Gradient Flow Path Length”
Fusheng Liu, Haizhao Yang, Soufiane Hayou and Qianxiao Li · 2022
Later among the works it cites.
“Feature Learning and Signal Propagation in Deep Neural Networks”
Yizhang Lou, Chris Mingard and Soufiane Hayou · 2022
Later among the works it cites.
“Scaling ResNets in the Large-depth Regime”
Pierre Marion, Adeline Fermanian, Gérard Biau and Jean-Philippe Vert · 2022
Later among the works it cites.
“Analyzing Finite Neural Networks: Can We Trust Neural Tangent Kernel Theory?”
Mariia Seleznova and Gitta Kutyniok · 2022
Later among the works it cites.
“Gaussian Pre-Activations in Neural Networks: Myth or Reality?”
Pierre Wolinski and Julyan Arbel · 2022
Later among the works it cites.
Greg Yang et al · 2022
Later among the works it cites.
“Efficient Computation of Deep Nonlinear Infinite-Width Neural Networks that Learn Features”
Greg Yang, Michael Santacroce and Edward Hu · 2022
Later among the works it cites.
“Deep Learning without Shortcuts: Shaping the Kernel with Tailored Rectifiers”
Guodong Zhang, Aleksandar Botev and James Martens · 2022
Later among the works it cites.