Fetching the paper…
Reading the bibliography…
There is a growing amount of literature on the relationship between wide neural networks (NNs) and Gaussian processes (GPs), identifying an equivalence between the two for a variety of NN architectures.
Yang, G · 1902
Earlier work this paper cites.
Central limit theorems for interchangeable processes
Blum, J. R., Chernoff, H., Rosenblatt, M., and Teicher, H · 1958
Earlier work this paper cites.
Probability and Measure
Billingsley, P · 1986
Earlier work this paper cites.
Bayesian Learning for Neural Networks
Neal, R. M · 1996
Earlier work this paper cites.
Efficient backprop
LeCun, Y., Bottou, L., Orr, G. B., and Müller, K.-R · 1998
Earlier work this paper cites.
Real Analysis and Probability
Dudley, R. M · 2002
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Enhanced convolutional neural kernels, 2020
Yu, D., Wang, R., Li, Z., Hu, W., Salakhutdinov, R., Arora, S., and Du, S. S · 2009
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C. D · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Convolutional sequence to sequence learning
Gehring, J., Auli, M., Grangier, D., Yarats, D., and Dauphin, Y. N · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Residual attention network for image classification
Wang, F., Jiang, M., Qian, C., Yang, S., Li, C., Zhang, H., Wang, X., and Tang, X · 2017
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., and Wanderman-Milne, S · 2018
Cited alongside, same era.
Aˆ 2-nets: Double attention networks
Chen, Y., Kalantidis, Y., Li, J., Yan, S., and Feng, J · 2018
Cited alongside, same era.
Squeeze-and-excitation networks
Hu, J., Shen, L., and Sun, G · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z., Li, Y., and Song, Z · 2019
Later among the works it cites.
On exact computation with an infinitely wide neural net
Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R. R., and Wang, R · 2019
Later among the works it cites.
Attention augmented convolutional networks
Bello, I., Zoph, B., Vaswani, A., Shlens, J., and Le, Q. V · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Later among the works it cites.
Gradient descent finds global minima of deep neural networks
Du, S. S., Lee, J. D., Li, H., Wang, L., and Zhai, X · 2019
Later among the works it cites.
Deep convolutional networks as shallow gaussian processes
Garriga-Alonso, A., Rasmussen, C. E., and Aitchison, L · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep neural networks as gaussian processes
Lee, J., Sohl-dickstein, J., Pennington, J., Novak, R., Schoenholz, S., and Bahri, Y · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y. and Liang, Y · 2018
Cited alongside, same era.
Gaussian process behaviour in wide deep neural networks
Matthews, A. G., Rowland, M., Hron, J., Turner, R. E., and Ghahramani, Z · 2018
Cited alongside, same era.
Deep contextualized word representations
Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L · 2018
Cited alongside, same era.
Self-attention with relative position representations
Shaw, P., Uszkoreit, J., and Vaswani, A · 2018
Cited alongside, same era.
Non-local neural networks
Wang, X., Girshick, R., Gupta, A., and He, K · 2018
Cited alongside, same era.
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J., Xiao, L., Schoenholz, S., Bahri, Y., Novak, R., Sohl-Dickstein, J., and Pennington, J · 2019
Later among the works it cites.
Bayesian deep convolutional networks with many channels are gaussian processes
Novak, R., Xiao, L., Bahri, Y., Lee, J., Yang, G., Abolafia, D. A., Pennington, J., and Sohl-dickstein, J · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Later among the works it cites.
Stand-alone self-attention in vision models
Ramachandran, P., Parmar, N., Vaswani, A., Bello, I., Levskaya, A., and Shlens, J · 2019
Later among the works it cites.
On the relationship between self-attention and convolutional layers
Cordonnier, J.-B., Loukas, A., and Jaggi, M · 2020
Closest in time.
Neural tangents: Fast and easy infinite neural networks in python
Novak, R., Xiao, L., Hron, J., Lee, J., Alemi, A. A., Sohl-Dickstein, J., and Schoenholz, S. S · 2020
Closest in time.
Sentiment classification using document embeddings trained with cosine similarity
Thongtan, T. and Phienthrakul, T · 2057
Closest in time.