Fetching the paper…
Reading the bibliography…
How does the geometric representation of a dataset change after the application of each randomly initialized layer of a neural network? The celebrated Johnson--Lindenstrauss lemma answers this question for linear fully-connected neural networks (FNNs), stating that the geometry is essentially preserved.
Extensions of Lipschitz mappings into a Hilbert space
William B Johnson and Joram Lindenstrauss · 1984
Earlier work this paper cites.
An elementary proof of the Johnson–Lindenstrauss lemma
Sanjoy Dasgupta and Anupam Gupta · 1999
Earlier work this paper cites.
Kernel methods for deep learning
Youngmin Cho and Lawrence Saul · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural network
Andrew M. Saxe, James L. Mcclelland, and Surya Ganguli · 2014
Earlier work this paper cites.
Suprema of chaos processes and the restricted isometry property
Felix Krahmer, Shahar Mendelson, and Holger Rauhut · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Deep neural networks with random Gaussian weights: A universal classification strategy?
Raja Giryes, Guillermo Sapiro, and Alex M Bronstein · 2016
Earlier work this paper cites.
Toward deeper understanding of neural networks: the power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Cited alongside, same era.
Exponential expressivity in deep neural networks through transient chaos
Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli · 2016
Cited alongside, same era.
Deep information propagation
Samuel S. Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein · 2017
Cited alongside, same era.
Mean field residual networks: On the edge of chaos
Ge Yang and Samuel Schoenholz · 2017
Cited alongside, same era.
Resurrecting the sigmoid in deep learning through dynamical isometry: Theory and practice
Jeffrey Pennington, Samuel S Schoenholz, and Surya Ganguli · 2017
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 2019
Later among the works it cites.
On the inductive bias of neural tangent kernels
Alberto Bietti and Julien Mairal · 2019
Later among the works it cites.
Initialization of relus for dynamical isometry
Rebekka Burkholz and Alina Dubatovka · 2019
Later among the works it cites.
Dynamical isometry is achieved in residual networks in a universal way for any activation function
Wojciech Tarnowski, Piotr Warchoł, Stanisław Jastrzębski, Jacek Tabor, and Maciej Nowak · 2019
Later among the works it cites.
Fixup initialization: Residual learning without normalization
Hongyi Zhang, Yann N. Dauphin, and Tengyu Ma · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
How to start training: The effect of initialization and architecture
Boris Hanin and David Rolnick · 2018
Cited alongside, same era.
Dynamical isometry and a mean field theory of cnns: How to train 10,000-layer vanilla convolutional neural networks
Lechao Xiao, Yasaman Bahri, Jascha Sohl-Dickstein, Samuel Schoenholz, and Jeffrey Pennington · 2018
Cited alongside, same era.
High-Dimensional Probability: An Introduction with Applications in Data Science
Roman Vershynin · 2018
Cited alongside, same era.
Alberto Bietti · 2021
Closest in time.
A review on weight initialization strategies for neural networks
Meenal V Narkhede, Prashant P Bartakke, and Mukul S Sutaone · 2021
Closest in time.
Deep networks and the multiple manifold problem
Sam Buchanan, Dar Gilboa, and John Wright · 2021
Closest in time.