Fetching the paper…
Reading the bibliography…
The performance of deep network learning strongly depends on the choice of the non-linear activation function associated with each neuron.
Multilayer feedforward networks are universal approximators
K. Hornik, M. B. Stinchcombe, and H. White · 1989
Earlier work this paper cites.
Multilayer feedforward networks with a nonpolynomial activation function can approximate any function
M. Leshno, V. Y. Lin, A. Pinkus, and S. Schocken · 1993
Earlier work this paper cites.
Padé approximations
C. Brezinski and J. Van Iseghem · 1994
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
N. Qian · 1999
Earlier work this paper cites.
Rate-coded restricted boltzmann machines for face recognition
Y. W. Teh and G. E. Hinton · 2001
Earlier work this paper cites.
Scalable parallel programming
J. Nickolls, I. Buck, and M. Garland · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky, G. Hinton, et al · 2009
Earlier work this paper cites.
Convolutional deep belief networks on cifar-10
A. Krizhevsky and G. Hinton · 2010
Earlier work this paper cites.
Mnist handwritten digit database. at&t labs, 2010
Y. LeCun, C. Cortes, and C. Burges · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
V. Nair and G. E. Hinton · 2010
Earlier work this paper cites.
Deep sparse rectifier neural networks
X. Glorot, A. Bordes, and Y. Bengio · 2011
Earlier work this paper cites.
Approximation Theory and Approximation Practice
L. N. Trefethen · 2012
Earlier work this paper cites.
Maxout networks
I. J. Goodfellow, D. Warde-Farley, M. Mirza, A. C. Courville, and Y. Bengio · 2013
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
A. L. Maas, A. Y. Hannun, and A. Y. Ng · 2013
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Cited alongside, same era.
Learning activation functions to improve deep neural networks
F. Agostinelli, M. D. Hoffman, P. J. Sadowski, and P. Baldi · 2015
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. S. Bernstein, A. C. Berg, and F. Li · 2015
Hyperactivations for activation function exploration
C. J. Vercellino and W. Y. Wang · 2017
Later among the works it cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
H. Xiao, K. Rasul, and R. Vollgraf · 2017
Later among the works it cites.
Improving deep convolutional neural networks with mixed maxout units
H.-Z. Zhao, F.-X. Liu, and L.-Y. Li · 2017
Later among the works it cites.
Understanding deep neural networks with rectified linear units
R. Arora, A. Basu, P. Mianjy, and A. Mukherjee · 2018
Later among the works it cites.
Rational neural networks for approximating graph convolution operator on jump discontinuities
Z. Chen, F. Chen, R. Lai, X. Zhang, and C.-T. Lu · 2018
Later among the works it cites.
Hardware implementation of hyperbolic tangent and sigmoid activation functions
Z. Hajduk · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Empirical evaluation of rectified activations in convolutional network
B. Xu, N. Wang, T. Chen, and M. Li · 2015
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (elus)
D. Clevert, T. Unterthiner, and S. Hochreiter · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Deep learning with s-shaped rectified linear activation units
X. Jin, C. Xu, J. Feng, Y. Wei, J. Xiong, and S. Yan · 2016
Cited alongside, same era.
Continuously differentiable exponential linear units
J. T. Barron · 2017
Cited alongside, same era.
Later among the works it cites.
Learning combinations of activation functions
F. Manessi and A. Rozza · 2018
Later among the works it cites.
Searching for activation functions
P. Ramachandran, B. Zoph, and Q. V. Le · 2018
Later among the works it cites.
Mobilenetv2: Inverted residuals and linear bottlenecks
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen · 2018
Later among the works it cites.
Provable robustness of relu networks via maximization of linear regions
F. Croce, M. Andriushchenko, and M. Hein · 2019
Closest in time.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
J. Frankle and M. Carbin · 2019
Closest in time.
Learning activation functions: A new paradigm of understanding neural networks
M. Goyal, R. Goyal, and B. Lall · 2019
Closest in time.
Universal approximation with deep narrow networks
P. Kidger and T. Lyons · 2019
Closest in time.