Fetching the paper…
Reading the bibliography…
We propose the Gaussian Error Linear Unit (GELU), a high-performing neural network activation function.
A logical calculus of the ideas immanent in nervous activity
Warren S. McCulloch and Walter Pitts · 1943
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
John Hopfield · 1982
Earlier work this paper cites.
Learning Multiple Layers of Features from Tiny Images, 2009
Alex Krizhevsky · 2009
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E. Hinton · 2010
Earlier work this paper cites.
Part-of-Speech Tagging for Twitter: Annotation, Features, and Experiments
Kevin Gimpel, Nathan Schneider, Brendan O ′ · 2011
Earlier work this paper cites.
Acoustic modeling using deep belief networks
Abdelrahman Mohamed, George E. Dahl, and Geoffrey E. Hinton · 2012
Earlier work this paper cites.
Adaptive dropout for training deep neural networks
Jimmy Ba and Brendan Frey · 2013
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Andrew L. Maas, Awni Y. Hannun, , and Andrew Y. Ng · 2013
Earlier work this paper cites.
Improved part-of-speech tagging for online conversational text with word clusters
Olutobi Owoputi, Brendan O’Connor, Chris Dyer, Kevin Gimpel, Nathan Schneider, and Noah A. Smith · 2013
Cited alongside, same era.
Improving neural networks with dropout
Nitish Srivastava · 2013
Cited alongside, same era.
Learning with pseudo-ensembles
Philip Bachman, Ouais Alsharif, and Doina Precup · 2014
Cited alongside, same era.
A simple approximation to the area under standard normal curve
Amit Choudhury · 2014
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M. Saxe, James L. McClelland, and Surya Ganguli · 2014
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey E. Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Fast and accurate deep network learning by exponential linear units (ELUs)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2016
Closest in time.
Adjusting for dropout variance in batch normalization and weight initialization
Dan Hendrycks and Kevin Gimpel · 2016
Closest in time.
Zoneout: Regularizing RNNs by randomly preserving hidden activations
David Krueger, Tegan Maharaj, János Kramár, Mohammad Pezeshki, Nicolas Ballas, Nan Rosemary Ke1, Anirudh Goyal, Yoshua Bengio, Hugo Larochelle, Aaron Courville, and Chris Pal · 2016
Closest in time.
SGDR: Stochastic gradient descent with restarts
Ilya Loshchilov and Frank Hutter · 2016
Closest in time.
All you need is a good init
Dmytro Mishkin and Jiri Matas · 2016
Closest in time.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Diederik P. Kingma · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Natural neural networks
Guillaume Desjardins, Karen Simonyan, Razvan Pascanu, and Koray Kavukcuoglu · 2015
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Diederik Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Closest in time.
Deep residual networks with exponential linear unit
Anish Shah, Sameer Shinde, Eashan Kadam, Hena Shah, and Sandip Shingade · 2016
Closest in time.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Closest in time.