Fetching the paper…
Reading the bibliography…
Yang (2020a) recently showed that the Neural Tangent Kernel (NTK) at initialization has an infinite-width limit for a large class of architectures including modern staples such as ResNet and Transformers.
Yang, G · 1902
Earlier work this paper cites.
Yang, G · 1910
Earlier work this paper cites.
Cognitron: A self-organizing multilayered neural network
Fukushima, K · 1975
Earlier work this paper cites.
Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position
Fukushima, K · 1980
Earlier work this paper cites.
Learning internal representations by error propagation
Rumelhart, D., Hinton, G. E., and Williams, R. J · 1986
Earlier work this paper cites.
Bayesian learning for neural networks
Hinton, G. E. and Neal, R · 1995
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Lecun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Object recognition with gradient-based learning
Lecun, Y., Haffner, P., and Bengio, Y · 2000
Earlier work this paper cites.
Littwin, E., Myara, B., Sabah, S., Susskind, J., Zhai, S., and Golan, O · 2006
Earlier work this paper cites.
Tensor programs ii: Neural tangent kernel for any architecture
Yang, G · 2006
Earlier work this paper cites.
Continuous neural networks
Roux, N. L. and Bengio, Y · 2007
Earlier work this paper cites.
Tensor programs iii: Neural matrix laws
Yang, G · 2009
Earlier work this paper cites.
Spectral networks and locally connected networks on graphs
Bruna, J., Zaremba, W., Szlam, A., and LeCun, Y · 2014
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Cho, K., Merrienboer, B. V., Çaglar Gülçehre, Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Cited alongside, same era.
Convolutional networks on graphs for learning molecular fingerprints
Duvenaud, D., Maclaurin, D., Aguilera-Iparraguirre, J., Gómez-Bombarelli, R., Hirzel, T., Aspuru-Guzik, A., and Adams, R · 2015
Cited alongside, same era.
Steps toward deep kernel methods from infinite neural networks
Hazan, T. and Jaakkola, T · 2015
Cited alongside, same era.
Gradient Descent Provably Optimizes Over-parameterized Neural Networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Later among the works it cites.
Deep neural networks as gaussian processes
Lee, J., Bahri, Y., Novak, R., Schoenholz, S., Pennington, J., and Sohl-Dickstein, J · 2018
Later among the works it cites.
Gaussian process behaviour in wide deep neural networks
Matthews, A., Rowland, M., Hron, J., Turner, R., and Ghahramani, Z · 2018
Later among the works it cites.
Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Henaff, M., Bruna, J., and LeCun, Y · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Daniely, A., Frostig, R., and Singer, Y · 2016
Cited alongside, same era.
Convolutional neural networks on graphs with fast localized spectral filtering
Defferrard, M., Bresson, X., and Vandergheynst, P · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Densely connected convolutional networks
Huang, G., Liu, Z., and Weinberger, K. Q · 2017
Cited alongside, same era.
Semi-supervised classification with graph convolutional networks
Kipf, T. and Welling, M · 2017
Cited alongside, same era.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
Zou, D., Cao, Y., Zhou, D., and Gu, Q · 2018
Later among the works it cites.
On exact computation with an infinitely wide neural net
Arora, S., Du, S., Hu, W., Li, Z., Salakhutdinov, R., and Wang, R · 2019
Later among the works it cites.
Graph neural tangent kernel: Fusing graph neural networks with graph kernels
Du, S., Hou, K., Póczos, B., Salakhutdinov, R., Wang, R., and Xu, K · 2019
Later among the works it cites.
Finite depth and width corrections to the neural tangent kernel
Hanin, B. and Nica, M · 2019
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J., Xiao, L., Schoenholz, S. S., Bahri, Y., Sohl-Dickstein, J., and Pennington, J · 2019
Later among the works it cites.
On random deep weight-tied autoencoders: Exact asymptotic analysis, phase transitions, and implications to training
Li, P. and Nguyen, P.-M · 2019
Later among the works it cites.
Bayesian deep convolutional networks with many channels are gaussian processes
Novak, R., Xiao, L., Bahri, Y., Lee, J., Yang, G., Hron, J., Abolafia, D., Pennington, J., and Sohl-Dickstein, J · 2019
Later among the works it cites.
The recurrent neural tangent kernel
Alemohammad, S., Wang, Z., Balestriero, R., and Baraniuk, R · 2020
Later among the works it cites.
Infinite attention: Nngp and ntk for deep attention networks
Hron, J., Bahri, Y., Sohl-Dickstein, J., and Novak, R · 2020
Later among the works it cites.
Feature learning in infinite-width neural networks
Yang, G. and Hu, E. J · 2020
Later among the works it cites.