Fetching the paper…
Reading the bibliography…
We revisit the initialization of deep residual networks (ResNets) by introducing a novel analytical tool in free probability to the community of deep learning.
V.A.Marchenko and L.A.Pastur, “Distribution of eigenvalues for some sets of random matrices,” Mathematics of the Ussr-Sbornik , vol. 1, no. 1, p. 507 & ndash;536, 1967
1967
Earlier work this paper cites.
V. Plerou, P. Gopikrishnan, B. Rosenow, L. A. N. Amaral, and H. E. Stanley, “Universal and nonuniversal properties of cross correlations in financial time series,” Physical Review Letters , vol. 83, no. 7, pp. 1471–1474, 1999
1999
Earlier work this paper cites.
A. Nica and R. Speicher, “Lectures on the combinatorics of free probability,” Cambridge Uk , 2006
2006
Earlier work this paper cites.
R. Collobert and J. Weston, “A unified architecture for natural language processing: Deep neural networks with multitask learning,” in Proceedings of the 25th international conference on Machine learning . ACM, 2008, pp. 160–167
2008
Earlier work this paper cites.
G. Ferraro, Lagrange inversion theorem . Springer New York, 2008
2008
Earlier work this paper cites.
A. Krizhevsky, “Learning multiple layers of features from tiny images,” 2009
2009
Earlier work this paper cites.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in Proceedings of the thirteenth international conference on artificial intelligence and statistics , 2010, pp. 249–256
2010
Earlier work this paper cites.
J. W. Silverstein, Spectral analysis of large dimensional random matrices . Science Press, 2010
2010
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in neural information processing systems , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
A.-r. Mohamed, G. E. Dahl, and G. Hinton, “Acoustic modeling using deep belief networks,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 20, no. 1, pp. 14–22, 2012
2012
Cited alongside, same era.
T. Tao, Topics in Random Matrix Theory . American Mathematical Society, 2012
2012
Cited alongside, same era.
B. Cakmak, “Non-hermitian random matrix theory for mimo channels,” Institutt for Elektronikk Og Telekommunikasjon , 2012
2012
Cited alongside, same era.
A. M. Saxe, J. L. Mcclelland, and S. Ganguli, “Exact solutions to the nonlinear dynamics of learning in deep linear neural networks,” Computer Science , 2013
2013
Cited alongside, same era.
G. Akemann, J. R. Ipsen, and M. Kieburg, “Products of rectangular random matrices: singular values and progressive scattering,” Phys Rev E Stat Nonlin Soft Matter Phys , vol. 88, no. 5, p. 052118, 2013
——, “Identity mappings in deep residual networks,” European Conference on Computer Vision , pp. 630–645, 2016
2016
Later among the works it cites.
J. Pennington, S. Schoenholz, and S. Ganguli, “Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice,” in Advances in Neural Information Processing Systems 30 , 2017, pp. 4785–4795
2017
Later among the works it cites.
M. Hardt and T. Ma, “Identity matters in deep learning,” International Conference on Learning Representations , 2017
2017
Later among the works it cites.
G. Yang and S. Schoenholz, “Mean field residual networks: On the edge of chaos,” in Advances in Neural Information Processing Systems 30 , 2017, pp. 7103–7114
2017
Later among the works it cites.
S. S. Schoenholz, J. Gilmer, S. Ganguli, and J. Sohldickstein, “Deep information propagation,” International Conference on Learning Representations , 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2013
Cited alongside, same era.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” Computer Science , 2014
2014
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,” International Conference on Computer Vision , pp. 1026–1034, 2015
2015
Cited alongside, same era.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in Proceedings of the 32nd international conference on Machine learning , 2015, pp. 448–456
2015
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Computer Vision and Pattern Recognition , 2016, pp. 770–778
2016
Cited alongside, same era.
Cited in the paper.
2017
Later among the works it cites.
J. Pennington, S. S. Schoenholz, and S. Ganguli, “The emergence of spectral universality in deep networks,” International Conference on Artificial Intelligence and Statistics , pp. 1924–1932, 2018
2018
Closest in time.
L. Xiao, Y. Bahri, J. Sohldickstein, S. S. Schoenholz, and J. Pennington, “Dynamical isometry and a mean field theory of cnns: How to train 10,000-layer vanilla convolutional neural networks,” in Proceedings of the 35th international conference on Machine learning , 2018, pp. 5389–5398
2018
Closest in time.
Z. Liao and R. Couillet, “On the spectrum of random features maps of high dimensional data,” International Conference on Machine Learning , pp. 3063–3071, 2018
2018
Closest in time.