Fetching the paper…
Reading the bibliography…
We introduce the "exponential linear unit" (ELU) which speeds up learning in deep neural networks and leads to higher classification accuracies.
Eigenvalues of covariance matrices: Application to neural-network learning
LeCun, Y., Kanter, I., and Solla, S. A · 1991
Earlier work this paper cites.
Iterative weighted least squares algorithms for neural networks classifiers
Kurita, T · 1993
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S.-I · 1998
Earlier work this paper cites.
The vanishing gradient problem during learning recurrent neural nets and problem solutions
Hochreiter, S · 1998
Earlier work this paper cites.
Efficient backprop
LeCun, Y., Bottou, L., Orr, G. B., and Müller, K.-R · 1998
Earlier work this paper cites.
Centering neural network gradient factor
Schraudolph, N. N · 1998
Earlier work this paper cites.
Complexity issues in natural gradient descent method for training multilayer perceptrons
Yang, H. H. and Amari, S.-I · 1998
Earlier work this paper cites.
Feature extraction through LOCOCODE
Hochreiter, S. and Schmidhuber, J · 1999
Earlier work this paper cites.
A Fast, Compact Approximation of the Exponential Function
Schraudolph, Nicol N · 1999
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies
Hochreiter, S., Bengio, Y., Frasconi, P., and Schmidhuber, J · 2001
Earlier work this paper cites.
Topmoumoute online natural gradient algorithm
LeRoux, N., Manzagol, P.-A., and Bengio, Y · 2008
Earlier work this paper cites.
Deep learning via Hessian-free optimization
Martens, J · 2010
Earlier work this paper cites.
Rectified linear units improve restricted Boltzmann machines
Nair, V. and Hinton, G. E · 2010
Cited alongside, same era.
Deep sparse rectifier neural networks
Glorot, X., Bordes, A., and Bengio, Y · 2011
Cited alongside, same era.
ImageNet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Cited alongside, same era.
Deep learning made easier by linear transformations in perceptrons
Raiko, T., Valpola, H., and LeCun, Y · 2012
Cited alongside, same era.
Krylov subspace descent for deep learning
Vinyals, O. and Povey, D · 2012
Cited alongside, same era.
Maxout networks
Goodfellow, I. J., Warde-Farley, D., Mirza, M., Courville, A., and Bengio, Y · 2013
Cited alongside, same era.
Striving for simplicity: The all convolutional net
Springenberg, Jost Tobias, Dosovitskiy, Alexey, Brox, Thomas, and Riedmiller, Martin A · 2014
Later among the works it cites.
Rectified factor networks
Clevert, D.-A., Unterthiner, T., Mayr, A., and Hochreiter, S · 2015
Closest in time.
Desjardins, G., Simonyan, K., Pascanu, R., and Kavukcuoglu, K · 2015
Closest in time.
Scaling up natural gradient by sparsely factorizing the inverse Fisher matrix
Grosse, R. and Salakhudinov, R · 2015
Closest in time.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lin, Min, Chen, Qiang, and Yan, Shuicheng · 2013
Cited alongside, same era.
Rectifier nonlinearities improve neural network acoustic models
Maas, A. L., Hannun, A. Y., and Ng, A. Y · 2013
Cited alongside, same era.
Riemannian metrics for neural networks i: feedforward networks
Olivier, Y · 2013
Cited alongside, same era.
Graham, Benjamin · 2014
Cited alongside, same era.
Learning Semantic Image Representations at a Large Scale
Jia, Yangqing · 2014
Cited alongside, same era.
Revisiting natural gradient for deep networks
Pascanu, R. and Bengio, Y · 2014
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Closest in time.
Deeply-supervised nets
Lee, Chen-Yu, Xie, Saining, Gallagher, Patrick W., Zhang, Zhengyou, and Tu, Zhuowen · 2015
Closest in time.
DeepTox: Toxicity prediction using deep learning
Mayr, A., Klambauer, G., Unterthiner, T., and Hochreiter, S · 2015
Closest in time.
Srivastava, Rupesh Kumar, Greff, Klaus, and Schmidhuber, Jürgen · 2015
Closest in time.
Toxicity prediction using deep learning
Unterthiner, T., Mayr, A., Klambauer, G., and Hochreiter, S · 2015
Closest in time.
Empirical evaluation of rectified activations in convolutional network
Xu, B., Wang, N., Chen, T., and Li, M · 2015
Closest in time.