Fetching the paper…
Reading the bibliography…
Under mild conditions on the network initialization we derive a power series expansion for the Neural Tangent Kernel (NTK) of arbitrarily deep feedforward networks in the infinite width limit.
Generalization guarantees for neural networks via harnessing the low-rank structure of the Jacobian
Samet Oymak, Zalan Fabian, Mingchen Li, and Mahdi Soltanolkotabi · 1906
Earlier work this paper cites.
A fine-grained spectral perspective on neural networks, 2019
Greg Yang and Hadi Salman · 1907
Earlier work this paper cites.
Bemerkungen zur Theorie der beschränkten Bilinearformen mit unendlich vielen Veränderlichen
J. Schur · 1911
Earlier work this paper cites.
Das asymptotische Verteilungsgesetz der Eigenwerte linearer partieller Differentialgleichungen (mit einer Anwendung auf die Theorie der Hohlraumstrahlung)
Hermann Weyl · 1912
Earlier work this paper cites.
On Wallis’ formula
Donat K. Kazarinoff · 1956
Earlier work this paper cites.
Backpropagation can give rise to spurious local minima even for networks without hidden layers
Eduardo D. Sontag and Héctor J. Sussmann · 1989
Earlier work this paper cites.
Bayesian Learning for Neural Networks
Radford M. Neal · 1996
Earlier work this paper cites.
Real analysis: Modern techniques and their applications
G. B. Folland · 1999
Earlier work this paper cites.
Neural Network Learning - Theoretical Foundations
Martin Anthony and Peter L. Bartlett · 2002
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L. Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
Andrea Caponnetto and Ernesto De Vito · 2007
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Efficient BackProp , pp. 9–48
Yann A. LeCun, Léon Bottou, Genevieve B. Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices , pp. 210–268
Roman Vershynin · 2012
Earlier work this paper cites.
Analysis of Boolean functions
Ryan O’Donnell · 2014
Earlier work this paper cites.
Eigenvalues of dot-product kernels on the sphere
Douglas Azevedo and Valdir A Menegatto · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Earlier work this paper cites.
Dmytro Mishkin and Jiri Matas · 2016
Earlier work this paper cites.
Exponential expressivity in deep neural networks through transient chaos
Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli · 2016
Earlier work this paper cites.
Nonlinear random matrix theory for deep learning
Jeffrey Pennington and Pratik Worah · 2017
Earlier work this paper cites.
Deep information propagation
Samuel S. Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Diverse Neural Network Learns True Target Functions
Bo Xie, Yingyu Liang, and Le Song · 2017
Cited alongside, same era.
Gaussian process behaviour in wide deep neural networks
Alexander G. de G. Matthews, Jiri Hron, Mark Rowland, Richard E. Turner, and Zoubin Ghahramani · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Cited alongside, same era.
Deep neural networks as Gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S. Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Cited alongside, same era.
A random matrix approach to neural networks
Cosme Louart, Zhenyu Liao, and Romain Couillet · 2018
Cited alongside, same era.
The spectrum of the Fisher information matrix of a single-hidden-layer neural network
Jeffrey Pennington and Pratik Worah · 2018
Effect of activation functions on the training of overparametrized neural nets
Abhishek Panigrahi, Abhishek Shetty, and Navin Goyal · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D. Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Later among the works it cites.
Gradient descent optimizes over-parameterized deep ReLU networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2020
Later among the works it cites.
Deep equals shallow for ReLU networks in kernel regimes
Alberto Bietti and Francis Bach · 2021
Later among the works it cites.
Deep neural tangent kernel and laplace kernel have the same RKHS
Lin Chen and Sheng Xu · 2021
Later among the works it cites.
Generalization error rates in kernel regression: The crossover from the noiseless to noisy regime
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Russ R Salakhutdinov, and Ruosong Wang · 2019
Cited alongside, same era.
The convergence rate of neural networks for learned functions of different frequencies
Ronen Basri, David W. Jacobs, Yoni Kasten, and Shira Kritchman · 2019
Cited alongside, same era.
On the inductive bias of neural tangent kernels
Alberto Bietti and Julien Mairal · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Cited alongside, same era.
Bayesian deep convolutional networks with many channels are Gaussian processes
Roman Novak, Lechao Xiao, Yasaman Bahri, Jaehoon Lee, Greg Yang, Jiri Hron, Daniel A. Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2019
Cited alongside, same era.
Hugo Cui, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová · 2021
Later among the works it cites.
A general expression for Hermite expansions with applications
Tom Davis · 2021
Later among the works it cites.
On the proof of global convergence of gradient descent for deep relu networks with linear widths
Quynh Nguyen · 2021
Later among the works it cites.
Tight bounds on the smallest eigenvalue of the neural tangent kernel for deep ReLU networks
Quynh Nguyen, Marco Mondelli, and Guido Montúfar · 2021
Later among the works it cites.
A spectral analysis of dot-product kernels
Meyer Scetbon and Zaid Harchaoui · 2021
Later among the works it cites.
Explicit loss asymptotics in the gradient descent training of neural networks
Maksim Velikanov and Dmitry Yarotsky · 2021
Later among the works it cites.
Implicit bias of MSE gradient optimization in underparameterized neural networks
Benjamin Bowman and Guido Montúfar · 2022
Closest in time.
Spectral bias outside the training set for deep networks in the kernel regime
Benjamin Bowman and Guido Montufar · 2022
Closest in time.
TorchNTK: A library for calculation of neural tangent kernels of PyTorch models
Andrew Engel, Zhichao Wang, Anand Sarwate, Sutanay Choudhury, and Tony Chiang · 2022
Closest in time.
On the spectral bias of convolutional neural tangent and gaussian process kernels
Amnon Geifman, Meirav Galun, David Jacobs, and Ronen Basri · 2022
Closest in time.
Fast neural kernel embeddings for general activations
Insu Han, Amir Zandieh, Jaehoon Lee, Roman Novak, Lechao Xiao, and Amin Karbasi · 2022
Closest in time.
Dimensionality reduction, regularization, and generalization in overparameterized regressions
Ningyuan Teresa Huang, David W. Hogg, and Soledad Villar · 2022
Closest in time.
Learning curves for gaussian process regression with power-law priors and targets
Hui Jin, Pradeep Kr. Banerjee, and Guido Montúfar · 2022
Closest in time.
Caltech 101, Apr 2022
Li, Andreeto, Ranzato, and Perona · 2022
Closest in time.
Activation function design for deep networks: linearity and effective initialisation
M. Murray, V. Abrol, and J. Tanner · 2022
Closest in time.
Fast finite width neural tangent kernel
Roman Novak, Jascha Sohl-Dickstein, and Samuel S Schoenholz · 2022
Closest in time.
Reverse engineering the neural tangent kernel
James Benjamin Simon, Sajant Anand, and Mike Deweese · 2022
Closest in time.