Fetching the paper…
Reading the bibliography…
We prove that a randomly initialized neural network of *any architecture* has its Tangent Kernel (NTK) converge to a deterministic limit, as the network widths tend to infinity.
Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient Descent
Jaehoon Lee, Lechao Xiao, Samuel S. Schoenholz, Yasaman Bahri, Jascha Sohl-Dickstein, and Jeffrey Pennington · 1902
Earlier work this paper cites.
Greg Yang · 1902
Earlier work this paper cites.
A Mean Field Theory of Batch Normalization
Greg Yang, Jeffrey Pennington, Vinay Rao, Jascha Sohl-Dickstein, and Samuel S. Schoenholz · 1902
Earlier work this paper cites.
Towards Characterizing Divergence in Deep Q-Learning
Joshua Achiam, Ethan Knight, and Pieter Abbeel · 1903
Earlier work this paper cites.
On Exact Computation with an Infinitely Wide Neural Net
Sanjeev Arora, Simon S. Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 1904
Earlier work this paper cites.
Cognitron: A self-organizing multilayered neural network
Kunihiko Fukushima · 1975
Earlier work this paper cites.
Neocognitron: A self-organizing neural network model for a mechanism of visual pattern recognition
Kunihiko Fukushima and Sei Miyake · 1982
Earlier work this paper cites.
Learning internal representations by error propagation
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams · 1985
Earlier work this paper cites.
BAYESIAN LEARNING FOR NEURAL NETWORKS
Radford M Neal · 1995
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Computing with Infinite Networks
Christopher K I Williams · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Object recognition with gradient-based learning
Yann LeCun, Patrick Haffner, Léon Bottou, and Yoshua Bengio · 1999
Earlier work this paper cites.
Continuous neural networks
Nicolas Le Roux and Yoshua Bengio · 2007
Earlier work this paper cites.
Kernel methods for deep learning
Youngmin Cho and Lawrence K. Saul · 2009
Earlier work this paper cites.
Spectral Networks and Locally Connected Networks on Graphs
Joan Bruna, Wojciech Zaremba, Arthur Szlam, and Yann LeCun · 2013
Earlier work this paper cites.
Neural Machine Translation by Jointly Learning to Align and Translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Generative Adversarial Networks
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Convolutional Networks on Graphs for Learning Molecular Fingerprints
David K Duvenaud, Dougal Maclaurin, Jorge Iparraguirre, Rafael Bombarell, Timothy Hirzel, Alan Aspuru-Guzik, and Ryan P Adams · 2015
Earlier work this paper cites.
Steps Toward Deep Kernel Methods from Infinite Neural Networks
Tamir Hazan and Tommi Jaakkola · 2015
Earlier work this paper cites.
Deep Convolutional Networks on Graph-Structured Data
Mikael Henaff, Joan Bruna, and Yann LeCun · 2015
Cited alongside, same era.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Gated Graph Sequence Neural Networks
Yujia Li, Daniel Tarlow, Marc Brockschmidt, and Richard Zemel · 2015
Cited alongside, same era.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Cited alongside, same era.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Cited alongside, same era.
Deep Neural Networks as Gaussian Processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Sam Schoenholz, Jeffrey Pennington, and Jascha Sohl-dickstein · 2018
Later among the works it cites.
Gaussian Process Behaviour in Wide Deep Neural Networks
Alexander G. de G. Matthews, Mark Rowland, Jiri Hron, Richard E. Turner, and Zoubin Ghahramani · 2018
Later among the works it cites.
Bayesian Deep Convolutional Networks with Many Channels are Gaussian Processes
Roman Novak, Lechao Xiao, Jaehoon Lee, Yasaman Bahri, Daniel A Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Later among the works it cites.
The Nonlinearity Coefficient - Predicting Overfitting in Deep Neural Networks
George Philipp and Jaime G. Carbonell · 2018
Later among the works it cites.
Dynamical Isometry and a Mean Field Theory of CNNs: How to Train 10,000-Layer Vanilla Convolutional Neural Networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Toward Deeper Understanding of Neural Networks: The Power of Initialization and a Dual View on Expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Cited alongside, same era.
Convolutional Neural Networks on Graphs with Fast Localized Spectral Filtering
Michaël Defferrard, Xavier Bresson, and Pierre Vandergheynst · 2016
Cited alongside, same era.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Densely Connected Convolutional Networks
Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q. Weinberger · 2016
Cited alongside, same era.
Semi-Supervised Classification with Graph Convolutional Networks
Thomas N. Kipf and Max Welling · 2016
Cited alongside, same era.
Exponential expressivity in deep neural networks through transient chaos
Ben Poole, Subhaneil Lahiri, Maithreyi Raghu, Jascha Sohl-Dickstein, and Surya Ganguli · 2016
Cited alongside, same era.
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
Jeffrey Pennington, Samuel Schoenholz, and Surya Ganguli · 2017
Cited alongside, same era.
Lechao Xiao, Yasaman Bahri, Jascha Sohl-Dickstein, Samuel Schoenholz, and Jeffrey Pennington · 2018
Later among the works it cites.
Deep mean field theory: Layerwise variance and width variation as methods to control gradient explosion, 2018
Greg Yang and Sam S. Schoenholz · 2018
Later among the works it cites.
Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Later among the works it cites.
The Convergence Rate of Neural Networks for Learned Functions of Different Frequencies
Ronen Basri, David Jacobs, Yoni Kasten, and Shira Kritchman · 2019
Later among the works it cites.
Neural temporal-difference and q-learning provably converge to global optima, 2019
Qi Cai, Zhuoran Yang, Jason D. Lee, and Zhaoran Wang · 2019
Later among the works it cites.
Graph neural tangent kernel: Fusing graph neural networks with graph kernels, 2019
Simon S. Du, Kangcheng Hou, Barnabás Póczos, Ruslan Salakhutdinov, Ruosong Wang, and Keyulu Xu · 2019
Later among the works it cites.
Asymptotics of wide networks from feynman diagrams, 2019
Ethan Dyer and Guy Gur-Ari · 2019
Later among the works it cites.
Linearized two-layers neural networks in high dimension
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Later among the works it cites.
Finite depth and width corrections to the neural tangent kernel, 2019
Boris Hanin and Mihai Nica · 2019
Later among the works it cites.
Dynamics of deep neural networks and neural tangent hierarchy, 2019
Jiaoyang Huang and Horng-Tzer Yau · 2019
Later among the works it cites.
A fine-grained spectral perspective on neural networks, 2019
Greg Yang and Hadi Salman · 2019
Later among the works it cites.
The recurrent neural tangent kernel, 2020
Sina Alemohammad, Zichao Wang, Randall Balestriero, and Richard Baraniuk · 2020
Closest in time.
Infinite attention: Nngp and ntk for deep attention networks, 2020
Jiri Hron, Yasaman Bahri, Jascha Sohl-Dickstein, and Roman Novak · 2020
Closest in time.
On the neural tangent kernel of deep networks with orthogonal initialization, 2020
Wei Huang, Weitao Du, and Richard Yi Da Xu · 2020
Closest in time.
Residual tangent kernels, 2020
Etai Littwin and Lior Wolf · 2020
Closest in time.
Tensor programs iii: Neural matrix laws
Greg Yang · 2020
Closest in time.