Fetching the paper…
Reading the bibliography…
The weight initialization and the activation function of deep neural networks have a crucial impact on the performance of the training procedure.
Bayesian learning for neural networks
R.M. Neal · 1995
Earlier work this paper cites.
Efficient backprop
Y. LeCun, L. Bottou, G. Orr, and K. Muller · 1998
Earlier work this paper cites.
Gradient flow in recurrent nets: The difficulty of learning longterm dependencies
J.F. Kolen and S.C. Kremer · 2001
Earlier work this paper cites.
Kernel methods for deep learning
Y. Cho and L.K. Saul · 2009
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
R. Pascanu, T. Mikolov, and Y. Bengio · 2013
Earlier work this paper cites.
On the number of linear regions of deep neural networks
G.F. Montufar, R. Pascanu, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Probabilistic backpropagation for scalable learning of bayesian neural networks
J. M. Hernandez-Lobato and R.P. Adams · 2015
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (elus)
D.A. Clevert, T. Unterthiner, and S. Hochreiter · 2016
Cited alongside, same era.
Bridging nonlinearities and stochastic regularizers with gaussian error linear units
D.. Hendrycks and K. Gimpel · 2016
Cited alongside, same era.
Exponential expressivity in deep neural networks through transient chaos
B. Poole, S. Lahiri, M. Raghu, J. Sohl-Dickstein, and S. Ganguli · 2016
Cited alongside, same era.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
Self-normalizing neural networks
G. Klambauer, T. Unterthiner, and A. Mayr · 2017
Later among the works it cites.
Searching for activation functions
P. Ramachandran, B. Zoph, and Q.V. Le · 2017
Later among the works it cites.
Deep information propagation
S.S. Schoenholz, J. Gilmer, S. Ganguli, and J. Sohl-Dickstein · 2017
Later among the works it cites.
Deep neural networks as gaussian processes
J. Lee, Y. Bahri, R. Novak, S.S. Schoenholz, J. Pennington, and J. Sohl-Dickstein · 2018
Closest in time.
Gaussian process behaviour in wide deep neural networks
A.G. Matthews, J. Hron, M. Rowland, R.E. Turner, and Z. Ghahramani · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Elfwing, E. Uchibe, and K. Doya · 2017
Cited alongside, same era.