Fetching the paper…
Reading the bibliography…
We present the surprising result that randomly initialized neural networks are good feature extractors in expectation.
On exact computation with an infinitely wide neural net
Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R., and Wang, R · 1904
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate of ( 1 / k 2 1/k^{2} )
Nesterov, Y. E · 1983
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
LeCun, Y., Boser, B., Denker, J. S., Henderson, D., Howard, R. E., Hubbard, W., and Jackel, L. D · 1989
Earlier work this paper cites.
On the limited memory BFGS method for large scale optimization
Liu, D. C. and Nocedal, J · 1989
Earlier work this paper cites.
Two algorithms for nearest-neighbor search in high dimensions
Kleinberg, J. M · 1997
Earlier work this paper cites.
An elementary proof of a theorem of johnson and lindenstrauss
Dasgupta, S. and Gupta, A · 2003
Earlier work this paper cites.
Random projection trees and low dimensional manifolds
Dasgupta, S. and Freund, Y · 2008
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
An analysis of deep neural network models for practical applications
Canziani, A., Paszke, A., and Culurciello, E · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Schoenholz, S. S., Gilmer, J., Ganguli, S., and Sohl-Dickstein, J · 2016
Cited alongside, same era.
Deep residual networks with exponential linear unit
Shah, A., Kadam, E., Shah, H., Shinde, S., and Shingade, S · 2016
Cited alongside, same era.
Instance normalization: The missing ingredient for fast stylization
Ulyanov, D., Vedaldi, A., and Lempitsky, V · 2016
Cited alongside, same era.
TriMap: Large-scale Dimensionality Reduction Using Triplets
Amid, E. and Warmuth, M. K · 2019
Later among the works it cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Cao, Y. and Gu, Q · 2019
Later among the works it cites.
Gradient descent finds global minima of deep neural networks
Du, S., Lee, J., Li, H., Wang, L., and Zhai, X · 2019
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J., Xiao, L., Schoenholz, S., Bahri, Y., Novak, R., Sohl-Dickstein, J., and Pennington, J · 2019
Later among the works it cites.
Fixup initialization: Residual learning without normalization
Zhang, H., Dauphin, Y. N., and Ma, T · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lee, J., Bahri, Y., Novak, R., Schoenholz, S. S., Pennington, J., and Sohl-Dickstein, J · 2017
Cited alongside, same era.
A gaussian process perspective on convolutional neural networks
Borovykh, A · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2018
Cited alongside, same era.
Deep convolutional networks as shallow Gaussian processes
Garriga-Alonso, A., Rasmussen, C. E., and Aitchison, L · 2018
Cited alongside, same era.
Decoupling Backpropagation using constrained optimization methods
Gotmare, A. D., Thomas, V., Brea, J., and Jaggi, M · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y. and Liang, Y · 2018
Cited alongside, same era.
Domingos, P · 2020
Later among the works it cites.
Rigging the lottery: Making all tickets winners
Evci, U., Gale, T., Menick, J., Castro, P. S., and Elsen, E · 2020
Later among the works it cites.
Training batchnorm and only batchnorm: On the expressive power of random features in CNNs
Frankle, J., Schwab, D. J., and Morcos, A. S · 2020
Later among the works it cites.
Training deep neural networks without batch normalization
Gaur, D., Folz, J., and Dengel, A · 2020
Later among the works it cites.
Infinite attention: NNGP and NTK for deep attention networks
Hron, J., Bahri, Y., Sohl-Dickstein, J., and Novak, R · 2020
Later among the works it cites.
What’s hidden in a randomly weighted neural network?
Ramanujan, V., Wortsman, M., Kembhavi, A., Farhadi, A., and Rastegari, M · 2020
Later among the works it cites.
Rezero is all you need: Fast convergence at large depth
Bachlechner, T., Majumder, B. P., Mao, H., Cottrell, G., and McAuley, J · 2021
Later among the works it cites.
On the equivalence between neural network and support vector machine
Chen, Y., Huang, W., Nguyen, L., and Weng, T.-W · 2021
Later among the works it cites.
Martens, J., Ballard, A., Desjardins, G., Swirszcz, G., Dalibard, V., Sohl-Dickstein, J., and Schoenholz, S. S · 2021
Later among the works it cites.