Fetching the paper…
Reading the bibliography…
Deep learning researchers have a keen interest in proposing two new novel activation functions which can boost network performance.
Backpropagation applied to handwritten zip code recognition
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel · 1989
Earlier work this paper cites.
Scalable parallel programming
J. Nickolls, I. Buck, M. Garland, and K. Skadron · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E. Hinton · 2010
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
Mark Everingham, Luc Gool, Christopher K. Williams, John Winn, and Andrew Zisserman · 2010
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Andrew L. Maas, Awni Y. Hannun, and Andrew Y. Ng · 2013
Earlier work this paper cites.
Maxout networks, 2013
Ian J. Goodfellow, David Warde-Farley, Mehdi Mirza, Aaron Courville, and Yoshua Bengio · 2013
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Improving deep neural networks using softplus units
Hao Zheng, Zhanlei Yang, Wenju Liu, Jizhong Liang, and Yanpeng Li · 2015
Earlier work this paper cites.
Empirical evaluation of rectified activations in convolutional network, 2015
Bing Xu, Naiyan Wang, Tianqi Chen, and Mu Li · 2015
Earlier work this paper cites.
Deep residual learning for image recognition, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition, 2015
Karen Simonyan and Andrew Zisserman · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift, 2015
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
U-net: Convolutional networks for biomedical image segmentation, 2015
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (elus), 2016
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2016
Cited alongside, same era.
Identity mappings in deep residual networks, 2016
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Ssd: Single shot multibox detector
Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C. Berg · 2016
Pytorch: An imperative style, high-performance deep learning library, 2019
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Later among the works it cites.
Mobilenetv2: Inverted residuals and linear bottlenecks, 2019
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen · 2019
Later among the works it cites.
Gaussian error linear units (gelus), 2020
Dan Hendrycks and Kevin Gimpel · 2020
Later among the works it cites.
Mish: A self regularized non-monotonic activation function, 2020
Diganta Misra · 2020
Later among the works it cites.
Padé activation units: End-to-end learning of flexible activation functions in deep networks, 2020
Alejandro Molina, Patrick Schramowski, and Kristian Kersting · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The cityscapes dataset for semantic urban scene understanding, 2016
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele · 2016
Cited alongside, same era.
Searching for activation functions, 2017
Prajit Ramachandran, Barret Zoph, and Quoc V. Le · 2017
Cited alongside, same era.
Shufflenet v2: Practical guidelines for efficient cnn architecture design, 2018
Ningning Ma, Xiangyu Zhang, Hai-Tao Zheng, and Jian Sun · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Cited alongside, same era.
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Later among the works it cites.
Universal approximation with deep narrow networks, 2020
Patrick Kidger and Terry Lyons · 2020
Later among the works it cites.
Efficientnet: Rethinking model scaling for convolutional neural networks, 2020
Mingxing Tan and Quoc V. Le · 2020
Later among the works it cites.
Erfact and pserf: Non-monotonic smooth trainable activation functions, 2021
Koushik Biswas, Sandeep Kumar, Shilpak Banerjee, and Ashish Kumar Pandey · 2021
Closest in time.
Orthogonal-padé activation functions: Trainable activation functions for smooth and faster convergence in deep networks, 2021
Koushik Biswas, Shilpak Banerjee, and Ashish Kumar Pandey · 2021
Closest in time.