Fetching the paper…
Reading the bibliography…
The activation function deployed in a deep neural network has great influence on the performance of the network at initialisation, which in turn has implications for training.
Learning representations by back-propagating errors
David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams · 1986
Earlier work this paper cites.
Bayesian Learning for Neural Networks
Radford M. Neal · 1996
Earlier work this paper cites.
On the lambertw function
R. M. Corless, G. H. Gonnet, D. E. G. Hare, D. J. Jeffrey, and D. E. Knuth · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
The vanishing gradient problem during learning recurrent neural nets and problem solutions
Sepp Hochreiter · 1998
Earlier work this paper cites.
Real analysis : modern techniques and their applications
G. B. Folland · 1999
Earlier work this paper cites.
An improved approximation for the gaussian q-function
G. K. Karagiannidis and A. S. Lioumpas · 2007
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M. Saxe, James L. McClelland, and Surya Ganguli · 2014
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (ELUs)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2015
Earlier work this paper cites.
Training very deep networks
Rupesh . Srivastava, Klaus Greff, and Jürgen Schmidhuber · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Exponential expressivity in deep neural networks through transient chaos
Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Cited alongside, same era.
Geometry of neural network loss surfaces via random matrix theory
Jeffrey Pennington and Yasaman Bahri · 2017
Later among the works it cites.
Searching for activation functions
Prajit Ramachandran, Barret Zoph, and Quoc V. Le · 2018
Later among the works it cites.
The emergence of spectral universality in deep networks
Jeffrey Pennington, Samuel Schoenholz, and Surya Ganguli · 2018
Later among the works it cites.
Gaussian process behaviour in wide deep neural networks
Alexander G. de G. Matthews, Jiri Hron, Mark Rowland, Richard E. Turner, and Zoubin Ghahramani · 2018
Later among the works it cites.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S. Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Later among the works it cites.
How to start training: The effect of initialization and architecture
Boris Hanin and David Rolnick · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ludovic Trottier, Philippe Giguère, and Brahim Chaib-draa · 2017
Cited alongside, same era.
Self-normalizing neural networks
Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter · 2017
Cited alongside, same era.
Deep information propagation
Samuel S. Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein · 2017
Cited alongside, same era.
Resurrecting the sigmoid in deep learning through dynamical isometry: Theory and practice
Jeffrey Pennington, Samuel S. Schoenholz, and Surya Ganguli · 2017
Cited alongside, same era.
Guide to Convolutional Neural Networks: A Practical Application to Traffic-Sign Detection and Classification
Hamed Aghdam and Elnaz Heravi · 2017
Cited alongside, same era.
Stable architectures for deep neural networks
Eldad Haber and Lars Ruthotto · 2017
Cited alongside, same era.
Later among the works it cites.
Which neural net architectures give rise to exploding and vanishing gradients?
Boris Hanin · 2018
Later among the works it cites.
On the impact of the activation function on deep neural networks training
Soufiane Hayou, Arnaud Doucet, and Judith Rousseau · 2019
Later among the works it cites.
The Random Matrix Theory of the Classical Compact Groups
Elizabeth S. Meckes · 2019
Later among the works it cites.
Tanhexp: A smooth activation function with high convergence speed for lightweight neural networks
Xinyu Liu and Xiaoguang Di · 2021
Closest in time.