Fetching the paper…
Reading the bibliography…
The choice of activation function in deep networks has a significant effect on the training dynamics and task performance.
“Batch renormalization: Towards reducing minibatch dependence in batch-normalized models,”
Sergey Ioffe, · 1953
Earlier work this paper cites.
“Representation of functions by superpositions of a step or sigmoid function and their applications to neural network theory,”
Yoshifusa Ito, · 1991
Earlier work this paper cites.
“An automated tanh-function method for finding solitary wave solutions to non-linear evolution equations,”
EJ Parkes and BR Duffy, · 1996
Earlier work this paper cites.
“Learning deep architectures for ai,”
Yoshua Bengio et al., · 2009
Earlier work this paper cites.
“Quadratic polynomials learn better image features,”
James Bergstra, Guillaume Desjardins, Pascal Lamblin, and Yoshua Bengio, · 2009
Earlier work this paper cites.
“Learning multiple layers of features from tiny images,”
Alex Krizhevsky, Geoffrey Hinton, et al., · 2009
Earlier work this paper cites.
“Rectified linear units improve restricted boltzmann machines,”
Vinod Nair and Geoffrey E Hinton, · 2010
Earlier work this paper cites.
“Understanding the difficulty of training deep feedforward neural networks,”
Xavier Glorot and Yoshua Bengio, · 2010
Earlier work this paper cites.
“Efficient backprop,”
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller, · 2012
Earlier work this paper cites.
“Rectifier nonlinearities improve neural network acoustic models,”
Andrew L Maas, Awni Y Hannun, and Andrew Y Ng, · 2013
Earlier work this paper cites.
Ian J Goodfellow, David Warde-Farley, Mehdi Mirza, Aaron Courville, and Yoshua Bengio, · 2013
Earlier work this paper cites.
“Even mirror fourier nonlinear filters,”
Alberto Carini and Giovanni L Sicuranza, · 2013
Earlier work this paper cites.
Min Lin, Qiang Chen, and Shuicheng Yan, · 2013
Earlier work this paper cites.
“Learning activation functions to improve deep neural networks,”
Forest Agostinelli, Matthew Hoffman, Peter Sadowski, and Pierre Baldi, · 2014
Earlier work this paper cites.
“Dropout: a simple way to prevent neural networks from overfitting,”
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Deep learning in neural networks: An overview,”
Jürgen Schmidhuber, · 2015
Cited alongside, same era.
“Deep learning,”
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton, · 2015
Cited alongside, same era.
“Fast and accurate deep network learning by exponential linear units (elus),”
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter, · 2015
Cited alongside, same era.
“Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2015
Cited alongside, same era.
“Empirical evaluation of rectified activations in convolutional network,”
Bing Xu, Naiyan Wang, Tianqi Chen, and Mu Li, · 2015
Cited alongside, same era.
“Parametric exponential linear unit for deep convolutional neural networks,”
Ludovic Trottier, Philippe Gigu, Brahim Chaib-draa, et al., · 2017
Later among the works it cites.
“Improving deep learning by inverse square root linear units (isrlus),”
Brad Carlile, Guy Delamarter, Paul Kinney, Akiko Marti, and Brian Whitney, · 2017
Later among the works it cites.
“P-telu: Parametric tan hyperbolic linear unit activation for deep neural networks,”
Rahul Duggal and Anubha Gupta, · 2017
Later among the works it cites.
“Deeparchitect: Automatically designing and training deep architectures,”
Renato Negrinho and Geoff Gordon, · 2017
Later among the works it cites.
“Mobilenets: Efficient convolutional neural networks for mobile vision applications,”
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam, · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sergey Ioffe and Christian Szegedy, · 2015
Cited alongside, same era.
“Going deeper with convolutions,”
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich, · 2015
Cited alongside, same era.
“Neural machine translation of rare words with subword units,”
Rico Sennrich, Barry Haddow, and Alexandra Birch, · 2015
Cited alongside, same era.
“Designing neural network architectures using reinforcement learning,”
Bowen Baker, Otkrist Gupta, Nikhil Naik, and Ramesh Raskar, · 2016
Cited alongside, same era.
“Noisy activation functions,”
Caglar Gulcehre, Marcin Moczulski, Misha Denil, and Yoshua Bengio, · 2016
Cited alongside, same era.
“Understanding and improving convolutional neural networks via concatenated rectified linear units,”
Wenling Shang, Kihyuk Sohn, Diogo Almeida, and Honglak Lee, · 2016
Cited alongside, same era.
“Deep learning with s-shaped rectified linear activation units,”
Xiaojie Jin, Chunyan Xu, Jiashi Feng, Yunchao Wei, Junjun Xiong, and Shuicheng Yan, · 2016
Cited alongside, same era.
Later among the works it cites.
“Densely connected convolutional networks,”
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger, · 2017
Later among the works it cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin, · 2017
Later among the works it cites.
“Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms,” 2017
Han Xiao, Kashif Rasul, and Roland Vollgraf, · 2017
Later among the works it cites.
“Frelu: Flexible rectified linear units for improving convolutional neural networks,”
Suo Qiu, Xiangmin Xu, and Bolun Cai, · 2018
Later among the works it cites.
“Improving deep neural network with multiple parametric exponential linear units,”
Yang Li, Chunxiao Fan, Yong Li, Qiong Wu, and Yue Ming, · 2018
Later among the works it cites.
“The quest for the golden activation function,”
Mina Basirat and Peter M Roth, · 2018
Later among the works it cites.
“Sigmoid-weighted linear units for neural network function approximation in reinforcement learning,”
Stefan Elfwing, Eiji Uchibe, and Kenji Doya, · 2018
Later among the works it cites.
“Invertible residual networks,”
Jens Behrmann, Will Grathwohl, Ricky TQ Chen, David Duvenaud, and Jörn-Henrik Jacobsen, · 2018
Later among the works it cites.
“Glow: Generative flow with invertible 1x1 convolutions,”
Durk P Kingma and Prafulla Dhariwal, · 2018
Later among the works it cites.
“Mobilenetv2: Inverted residuals and linear bottlenecks,”
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen, · 2018
Later among the works it cites.
“Mish: A self regularized non-monotonic neural activation function,”
Diganta Misra, · 2019
Later among the works it cites.