Fetching the paper…
Reading the bibliography…
Empirical neural tangent kernels (eNTKs) can provide a good understanding of a given network's representation: they are often far less expensive to compute and applicable more broadly than infinite width NTKs.
Smoothing noisy data with spline functions
Craven, P. and Wahba, G · 1978
Earlier work this paper cites.
Priors for infinite networks , pp. 29–53
Neal, R. M · 1996
Earlier work this paper cites.
Computing with infinite networks
Williams, C · 1996
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Kernels for vector-valued functions: A review
Álvarez, M. A., Rosasco, L., and Lawrence, N. D · 2012
Earlier work this paper cites.
Steps toward deep kernel methods from infinite neural networks
Hazan, T. and Jaakkola, T · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Wide residual networks
Zagoruyko, S. and Komodakis, N · 2016
Earlier work this paper cites.
FALKON: An optimal large scale kernel method
Rudi, A., Carratino, L., and Rosasco, L · 2017
Earlier work this paper cites.
Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms
Xiao, H., Rasul, K., and Vollgraf, R · 2017
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
Deep neural networks as gaussian processes
Lee, J., Bahri, Y., Novak, R., Schoenholz, S. S., Pennington, J., and Sohl-Dickstein, J · 2018
Earlier work this paper cites.
Gaussian process behaviour in wide deep neural networks
Matthews, A. G. d. G., Rowland, M., Hron, J., Turner, R. E., and Ghahramani, Z · 2018
Earlier work this paper cites.
High-dimensional probability: An introduction with applications in data science
Vershynin, R · 2018
Cited alongside, same era.
Dynamical isometry and a mean field theory of CNNs: How to train 10,000-layer vanilla convolutional neural networks
Xiao, L., Bahri, Y., Sohl-Dickstein, J., Schoenholz, S., and Pennington, J · 2018
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R., and Wang, R · 2019
Cited alongside, same era.
The benefits of over-parameterization at initialization in deep ReLU networks
Arpit, D. and Bengio, Y · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J., Xiao, L., Schoenholz, S., Bahri, Y., Novak, R., Sohl-Dickstein, J., and Pennington, J · 2019
Cited alongside, same era.
Fourier features let networks learn high frequency functions in low dimensional domains
Tancik, M., Srinivasan, P., Mildenhall, B., Fridovich-Keil, S., Raghavan, N., Singhal, U., Ramamoorthi, R., Barron, J., and Ng, R · 2020
Later among the works it cites.
Disentangling trainability and generalization in deep neural networks
Xiao, L., Pennington, J., and Schoenholz, S · 2020
Later among the works it cites.
Tensor Programs II: Neural tangent kernel for any architecture
Yang, G · 2020
Later among the works it cites.
When vision transformers outperform resnets without pre-training or strong data augmentations
Chen, X., Hsieh, C.-J., and Gong, B · 2021
Later among the works it cites.
Deep active learning by leveraging training dynamics
Wang, H., Huang, W., Margenot, A., Tong, H., and He, J · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bayesian deep convolutional networks with many channels are Gaussian processes
Novak, R., Xiao, L., Lee, J., Bahri, Y., Yang, G., Hron, J., Abolafia, D. A., Pennington, J., and Sohl-Dickstein, J · 2019
Cited alongside, same era.
The effect of network width on stochastic gradient descent and generalization: an empirical study
Park, D., Sohl-Dickstein, J., Le, Q., and Smith, S · 2019
Cited alongside, same era.
High-Dimensional Statistics: A Non-Asymptotic Viewpoint
Wainwright, M. J · 2019
Cited alongside, same era.
Wide feedforward or recurrent neural networks of any architecture are Gaussian processes
Yang, G · 2019
Cited alongside, same era.
Exploring the uncertainty properties of neural networks’ implicit priors in the infinite-width limit
Adlam, B., Lee, J., Xiao, L., Pennington, J., and Snoek, J · 2020
Cited alongside, same era.
Harnessing the power of infinitely wide deep nets on small-data tasks
Arora, S., Du, S. S., Li, Z., Salakhutdinov, R., Wang, R., and Yu, D · 2020
Cited alongside, same era.
Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the neural tangent kernel
Fort, S., Dziugaite, G. K., Paul, M., Kharaghani, S., Roy, D. M., and Ganguli, S · 2020
Cited alongside, same era.
Tensor Programs IIb: Architectural universality of neural tangent kernel training dynamics
Yang, G. and Littwin, E · 2021
Later among the works it cites.
Meta-learning with neural tangent kernels
Zhou, Y., Wang, Z., Xian, J., Chen, C., and Xu, J · 2021
Later among the works it cites.
Generalization through the lens of leave-one-out error
Bachmann, G., Hofmann, T., and Lucchi, A · 2022
Closest in time.
Note to self: Hanson–wright inequality, 2022
Epperly, E · 2022
Closest in time.
A neural tangent kernel perspective of GANs
Franceschi, J.-Y., de Bézenac, E., Ayed, I., Chen, M., Lamprier, S., and Gallinari, P · 2022
Closest in time.
Making look-ahead active learning strategies feasible with neural tangent kernels
Mohamadi, M. A., Bae, W., and Sutherland, D. J · 2022
Closest in time.
Fast finite width neural tangent kernel
Novak, R., Sohl-Dickstein, J., and Schoenholz, S. S · 2022
Closest in time.
More than a toy: Random matrix models predict how real-world neural representations generalize
Wei, A., Hu, W., and Steinhardt, J · 2022
Closest in time.
A framework and benchmark for deep batch active learning for regression
Holzmüller, D., Zaverkin, V., Kästner, J., and Steinwart, I · 2023
Closest in time.