Fetching the paper…
Reading the bibliography…
In this paper, a novel neural network activation function, called Symmetrical Gaussian Error Linear Unit (SGELU), is proposed to obtain high performance.
Lisht: Non-parametric linearly scaled hyperbolic tangent activation function for neural networks
Roy, S. K., Manna, S., Dubey, S. R., and Chaudhuri, B. B. (2019) · 1901
Earlier work this paper cites.
Handbook of methods of applied statistics
Brillinger, D. R., Chakravarti, I. M., Laha, R. G., and Roy, J · 1967
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Nair, V. and Hinton, G. E · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (elus)
Clevert, D.-A., Unterthiner, T., and Hochreiter, S · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
Gaussian error linear units (gelus)
Hendrycks, D. and Gimpel, K · 2016
Later among the works it cites.
Swish: a self-gated activation function
Ramachandran, P., Zoph, B., and Le, Q. V · 2017
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T., Simonyan, K., and Hassabis, D. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…