Fetching the paper…
Reading the bibliography…
Recent work has shown that tight concentration of the entire spectrum of singular values of a deep network's input-output Jacobian around one at initialization can speed up learning by orders of magnitude.
Free random variables
Dan V Voiculescu, Ken J Dykema, and Alexandru Nica · 1992
Earlier work this paper cites.
Multiplicative functions on the lattice of non-crossing partitions and free convolution
Roland Speicher · 1994
Earlier work this paper cites.
On the lambertw function
Robert M Corless, Gaston H Gonnet, David EG Hare, David J Jeffrey, and Donald E Knuth · 1996
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Topics in random matrix theory
Terence Tao · 2012
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M. Saxe, James L. McClelland, and Surya Ganguli · 2014
Cited alongside, same era.
Dmytro Mishkin and Jiri Matas · 2015
Cited alongside, same era.
Exponential expressivity in deep neural networks through transient chaos
B. Poole, S. Lahiri, M. Raghu, J. Sohl-Dickstein, and S. Ganguli · 2016
Cited alongside, same era.
Deep Information Propagation
S. S. Schoenholz, J. Gilmer, S. Ganguli, and J. Sohl-Dickstein · 2016
Cited alongside, same era.
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
Jeffrey Pennington, Samuel Schoenholz, and Surya Ganguli · 2017
Cited alongside, same era.
Searching for Activation Functions
P. Ramachandran, B. Zoph, and Q. V. Le
Cited in the paper.
Free probability and random matrices
James A Mingo and Roland Speicher · 2017
Later among the works it cites.
Geometry of neural network loss surfaces via random matrix theory
Jeffrey Pennington and Yasaman Bahri · 2017
Later among the works it cites.
On the generalization of the lambert �� function
István Mező and Árpád Baricz · 2017
Later among the works it cites.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
Stefan Elfwing, Eiji Uchibe, and Kenji Doya · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…