Fetching the paper…
Reading the bibliography…
The proper initialization of weights is crucial for the effective training and fast convergence of deep neural networks (DNNs).
Eigenvalues and condition numbers of random matrices
Alan Edelman · 1988
Earlier work this paper cites.
The spectral radii and norms of large dimensional non-central random matrices
Jack W Silverstein · 1994
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Products of random matrices: Dimension and growth in norm
Vladislav Kargin et al · 2010
Earlier work this paper cites.
MNIST handwritten digit database
Yann LeCun and Corinna Cortes · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
ClausIE: clause-based open information extraction
Luciano Del Corro and Rainer Gemulla · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean · 2013
Cited alongside, same era.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Cited alongside, same era.
AIDA-light: High-throughput named-entity disambiguation
Dat Ba Nguyen, Johannes Hoffart, Martin Theobald, and Gerhard Weikum · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi et al · 2015
Cited alongside, same era.
Probabilistic programming in python using pymc3
John Salvatier, Thomas V Wiecki, and Christopher Fonnesbeck · 2016
Later among the works it cites.
Revise saturated activation functions
Bing Xu, Ruitong Huang, and Mu Li · 2016
Later among the works it cites.
An exploration of word embedding initialization in deep-learning tasks
Tom Kocmi and Ondřej Bojar · 2017
Later among the works it cites.
How to start training: The effect of initialization and architecture
Boris Hanin and David Rolnick · 2018
Later among the works it cites.
Lipschitz regularity of deep neural networks: analysis and efficient estimation
Aladin Virmaux and Kevin Scaman · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
Adjusting for dropout variance in batch normalization and weight initialization
Dan Hendrycks and Kevin Gimpel · 2016
Cited alongside, same era.
Tensorflow github discussion: Hessian fails on fused ops
Tensorflow GitHub
Cited in the paper.
Tensorflow github discussion: High-order derivates and for loops
Tensorflow GitHub
Cited in the paper.
CIFAR-10 (Canadian Institute for Advanced Research)
Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton
Cited in the paper.
Ashish Agarwal · 2019
Later among the works it cites.
How to initialize your network? robust initialization for weightnorm & resnets
Devansh Arpit, Víctor Campos, and Yoshua Bengio · 2019
Later among the works it cites.
Scipy 1.0: Fundamental algorithms for scientific computing in python
Virtanen et al · 2020
Closest in time.