Fetching the paper…
Reading the bibliography…
The Neural Tangent Kernel (NTK), defined as $\Theta_\theta^f(x_1, x_2) = \left[\partial f(\theta, x_1)\big/\partial \theta\right] \left[\partial f(\theta, x_2)\big/\partial \theta\right]^T$ where $\left[\partial f(\theta, \cdot)\big/\partial \theta\right]$ is a neural network (NN) Jacobian, has emerged as a central object of study in deep learning.
Arora, S., Du, S. S., Hu, W., Li, Z., and Wang, R · 1901
Earlier work this paper cites.
Solution directe de l’équation séculaire et de quelques problèmes analogues transcendants
Müntz, H · 1913
Earlier work this paper cites.
Priors for infinite networks (tech. rep. no. crg-tr-94-1)
Neal, R. M · 1994
Earlier work this paper cites.
Optimal accumulation of jacobian matrices by elimination methods on the dual computational graph
Naumann, U · 2004
Earlier work this paper cites.
Evaluating Derivatives
Griewank, A. and Walther, A · 2008
Earlier work this paper cites.
Optimal jacobian accumulation is np-complete
Naumann, U · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Autograd: Effortless gradients in numpy
Maclaurin, D., Duvenaud, D., and Adams, R. P · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
Martens, J. and Grosse, R · 2015
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., et al · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Wide residual networks
Zagoruyko, S. and Komodakis, N · 2016
Earlier work this paper cites.
Neural architecture search with reinforcement learning, 2016
Zoph, B. and Le, Q. V · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Geometry of neural network loss surfaces via random matrix theory
Pennington, J. and Bahri, Y · 2017
Earlier work this paper cites.
Deep information propagation
Schoenholz, S. S., Gilmer, J., Ganguli, S., and Sohl-Dickstein, J · 2017
Earlier work this paper cites.
A gaussian process perspective on convolutional neural networks
Borovykh, A · 2018
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., and Wanderman-Milne, S · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
Deep neural networks as gaussian processes
Lee, J., Bahri, Y., Novak, R., Schoenholz, S., Pennington, J., and Sohl-dickstein, J · 2018
Earlier work this paper cites.
Gaussian process behaviour in wide deep neural networks
Matthews, A., Hron, J., Rowland, M., Turner, R. E., and Ghahramani, Z · 2018
Earlier work this paper cites.
Tadam: Task dependent adaptive metric for improved few-shot learning
Oreshkin, B. N., López, P. R., and Lacoste, A · 2018
Earlier work this paper cites.
Dynamical isometry and a mean field theory of CNNs: How to train 10,000-layer vanilla convolutional neural networks
Xiao, L., Bahri, Y., Sohl-Dickstein, J., Schoenholz, S., and Pennington, J · 2018
Cited alongside, same era.
Onnx: Open neural network exchange
Bai, J., Lu, F., Zhang, K., et al · 2019
Cited alongside, same era.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Belkin, M., Hsu, D., Ma, S., and Mandal, S · 2019
Cited alongside, same era.
Metainit: Initializing learning by learning to initialize
Dauphin, Y. N. and Schoenholz, S · 2019
Cited alongside, same era.
Graph neural tangent kernel: Fusing graph neural networks with graph kernels
Du, S. S., Hou, K., Salakhutdinov, R. R., Poczos, B., Wang, R., and Xu, K · 2019
Cited alongside, same era.
Deep convolutional networks as shallow gaussian processes
Infinitely wide graph convolutional networks: semi-supervised learning via gaussian processes
Hu, J., Shen, J., Yang, B., and Shao, L · 2020
Later among the works it cites.
Finite versus infinite neural networks: an empirical study
Lee, J., Schoenholz, S., Pennington, J., Adlam, B., Xiao, L., Novak, R., and Sohl-Dickstein, J · 2020
Later among the works it cites.
Dataset meta-learning from kernel ridge-regression
Nguyen, T., Chen, Z., and Lee, J · 2020
Later among the works it cites.
Neural tangents: Fast and easy infinite neural networks in python
Novak, R., Xiao, L., Hron, J., Lee, J., Alemi, A. A., Sohl-Dickstein, J., and Schoenholz, S. S · 2020
Later among the works it cites.
Towards nngp-guided neural architecture search
Park, D. S., Lee, J., Peng, D., Cao, Y., and Sohl-Dickstein, J · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Garriga-Alonso, A., Aitchison, L., and Rasmussen, C. E · 2019
Cited alongside, same era.
Approximate inference turns deep networks into gaussian processes
Khan, M. E. E., Immer, A., Abedi, E., and Korzepa, M · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J., Xiao, L., Schoenholz, S. S., Bahri, Y., Novak, R., Sohl-Dickstein, J., and Pennington, J · 2019
Cited alongside, same era.
Bayesian deep convolutional networks with many channels are gaussian processes
Novak, R., Xiao, L., Lee, J., Bahri, Y., Yang, G., Hron, J., Abolafia, D. A., Pennington, J., and Sohl-Dickstein, J · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Cited alongside, same era.
A jamming transition from under-to over-parametrization affects generalization in deep learning
Spigler, S., Geiger, M., d’Ascoli, S., Sagun, L., Biroli, G., and Wyart, M · 2019
Cited alongside, same era.
Yang, G · 2019
Cited alongside, same era.
Later among the works it cites.
Fourier features let networks learn high frequency functions in low dimensional domains
Tancik, M., Srinivasan, P. P., Mildenhall, B., Fridovich-Keil, S., Raghavan, N., Singhal, U., Ramamoorthi, R., Barron, J. T., and Ng, R · 2020
Later among the works it cites.
Disentangling trainability and generalization in deep learning
Xiao, L., Pennington, J., and Schoenholz, S. S · 2020
Later among the works it cites.
Non-Gaussian processes and neural networks at finite widths
Yaida, S · 2020
Later among the works it cites.
Tensor programs ii: Neural tangent kernel for any architecture
Yang, G · 2020
Later among the works it cites.
Explaining neural scaling laws
Bahri, Y., Dyer, E., Kaplan, J., Lee, J., and Sharma, U · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Later among the works it cites.
A neural tangent kernel perspective of gans
Franceschi, J.-Y., de Bézenac, E., Ayed, I., Chen, M., Lamprier, S., and Gallinari, P · 2021
Later among the works it cites.
Decomposing reverse-mode automatic differentiation
Frostig, R., Johnson, M. J., Maclaurin, D., Paszke, A., and Radul, A · 2021
Later among the works it cites.
Neural net training dynamics, January 2021
Grosse, R · 2021
Later among the works it cites.
functorch: Jax-like composable function transforms for pytorch
Horace He, R. Z · 2021
Later among the works it cites.
Dataset distillation with infinitely wide convolutional networks
Nguyen, T., Novak, R., Xiao, L., and Lee, J · 2021
Later among the works it cites.
How to train your vit? data, augmentation, and regularization in vision transformers
Steiner, A., Kolesnikov, A., Zhai, X., Wightman, R., Uszkoreit, J., and Beyer, L · 2021
Later among the works it cites.
Mlp-mixer: An all-mlp architecture for vision, 2021
Tolstikhin, I., Houlsby, N., Kolesnikov, A., Beyer, L., Zhai, X., Unterthiner, T., Yung, J., Steiner, A., Keysers, D., Uszkoreit, J., Lucic, M., and Dosovitskiy, A · 2021
Later among the works it cites.
Meta-learning with neural tangent kernels
Zhou, Y., Wang, Z., Xian, J., Chen, C., and Xu, J · 2021
Later among the works it cites.
Ukraine vectors by vecteezy
Arfian, W · 2022
Closest in time.
You only linearize once: Tangents transpose to gradients
Radul, A., Paszke, A., Frostig, R., Johnson, M., and Maclaurin, D · 2022
Closest in time.