Fetching the paper…
Reading the bibliography…
As a popular approach to modeling the dynamics of training overparametrized neural networks (NNs), the neural tangent kernels (NTK) are known to fall behind real-world NNs in generalization ability.
Linearized two-layers neural networks in high dimension
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 1904
Earlier work this paper cites.
Harnessing the power of infinitely wide deep nets on small-data tasks
Sanjeev Arora, Simon S Du, Zhiyuan Li, Ruslan Salakhutdinov, Ruosong Wang, and Dingli Yu · 1910
Earlier work this paper cites.
A class of statistics with asymptotically normal distribution
Wassily Hoeffding et al · 1948
Earlier work this paper cites.
Analytic structure of schläfli function
Kazuhiko Aomoto · 1977
Earlier work this paper cites.
Priors for infinite networks
Radford M Neal · 1996
Earlier work this paper cites.
Computing with infinite networks
Christopher KI Williams · 1997
Earlier work this paper cites.
The information bottleneck method
Naftali Tishby, Fernando C Pereira, and William Bialek · 2000
Earlier work this paper cites.
Asymptotic statistics , volume 3
Aad W Van der Vaart · 2000
Earlier work this paper cites.
On kernel-target alignment
Nello Cristianini, John Shawe-Taylor, Andre Elisseeff, and Jaz S Kandola · 2002
Earlier work this paper cites.
Learning the kernel matrix with semidefinite programming
Gert RG Lanckriet, Nello Cristianini, Peter Bartlett, Laurent El Ghaoui, and Michael I Jordan · 2004
Earlier work this paper cites.
Measuring solid angles beyond dimension three
Jason M Ribando · 2006
Earlier work this paper cites.
Continuous neural networks
Nicolas Le Roux and Yoshua Bengio · 2007
Earlier work this paper cites.
The fast johnson–lindenstrauss transform and approximate nearest neighbors
Nir Ailon and Bernard Chazelle · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A Krizhevsky · 2009
Earlier work this paper cites.
Large-margin classification in infinite neural networks
Youngmin Cho and Lawrence K Saul · 2010
Earlier work this paper cites.
Multiple kernel learning algorithms
Mehmet Gönen and Ethem Alpaydin · 2011
Earlier work this paper cites.
L2 regularization for learning kernels
Corinna Cortes, Mehryar Mohri, and Afshin Rostamizadeh · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Steps toward deep kernel methods from infinite neural networks
Tamir Hazan and Tommi Jaakkola · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky · 2015
Cited alongside, same era.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim · 2017
Cited alongside, same era.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger · 2017
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Yuan Cao and Quanquan Gu · 2019
Later among the works it cites.
How much over-parameterization is sufficient to learn deep relu networks?
Zixiang Chen, Yuan Cao, Difan Zou, and Quanquan Gu · 2019
Later among the works it cites.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Later among the works it cites.
Graph neural tangent kernel: Fusing graph neural networks with graph kernels
Simon S Du, Kangcheng Hou, Russ R Salakhutdinov, Barnabas Poczos, Ruosong Wang, and Keyulu Xu · 2019
Later among the works it cites.
Stiffness: A new perspective on generalization in neural networks
Stanislav Fort, Paweł Krzysztof Nowak, Stanislaw Jastrzebski, and Srini Narayanan · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2017
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2018
Cited alongside, same era.
A note on lazy training in supervised differentiable programming
Lenaic Chizat and Francis Bach · 2018
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Simon S Du, Jason D Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2018
Cited alongside, same era.
Deep convolutional networks as shallow gaussian processes
Adrià Garriga-Alonso, Carl Edward Rasmussen, and Laurence Aitchison · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Finite depth and width corrections to the neural tangent kernel
Boris Hanin and Mihai Nica · 2019
Later among the works it cites.
Dynamics of deep neural networks and neural tangent hierarchy
Jiaoyang Huang and Horng-Tzer Yau · 2019
Later among the works it cites.
Ziwei Ji and Matus Telgarsky · 2019
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Later among the works it cites.
Enhanced convolutional neural tangent kernels
Zhiyuan Li, Ruosong Wang, Dingli Yu, Simon S Du, Wei Hu, Ruslan Salakhutdinov, and Sanjeev Arora · 2019
Later among the works it cites.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
Colin Wei, Jason D Lee, Qiang Liu, and Tengyu Ma · 2019
Later among the works it cites.
Greg Yang · 2019
Later among the works it cites.
On the power and limitations of random features for understanding neural networks
Gilad Yehudai and Ohad Shamir · 2019
Later among the works it cites.
Taylorized training: Towards better approximation of neural network training at finite width
Yu Bai, Ben Krause, Huan Wang, Caiming Xiong, and Richard Socher · 2020
Closest in time.
A generalized neural tangent kernel analysis for two-layer neural networks
Zixiang Chen, Yuan Cao, Quanquan Gu, and Tong Zhang · 2020
Closest in time.
The local elasticity of neural networks
Hangfeng He and Weijie J. Su · 2020
Closest in time.
Neural tangents: Fast and easy infinite neural networks in python
Roman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee, Alexander A. Alemi, Jascha Sohl-Dickstein, and Samuel S. Schoenholz · 2020
Closest in time.
Gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2020
Closest in time.