Fetching the paper…
Reading the bibliography…
Deep learning at its core, contains functions that are composition of a linear transformation with a non-linear function known as activation function.
Digital selection and analogue amplification coexist in a cortex-inspired silicon circuit
Richard Hahnloser, Rahul Sarpeshkar, Misha Mahowald, Rodney Douglas, and H. Seung · 2000
Earlier work this paper cites.
Incorporating second-order functional knowledge for better option pricing
Charles Dugas, Yoshua Bengio, François Bélisle, Claude Nadeau, and René Garcia · 2000
Earlier work this paper cites.
What is the best multi-stage architecture for object recognition?
Kevin Jarrett, Koray Kavukcuoglu, Marc’Aurelio Ranzato, and Yann LeCun · 2009
Earlier work this paper cites.
Quadratic features and deep architectures for chunking
Joseph Turian, James Bergstra, and Yoshua Bengio · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E. Hinton · 2010
Earlier work this paper cites.
Mnist handwritten digit database
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Andrew L. Maas, Awni Y. Hannun, and Andrew Y. Ng · 2013
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus), 2015
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Empirical evaluation of rectified activations in convolutional network, 2015
Bing Xu, Naiyan Wang, Tianqi Chen, and Mu Li · 2015
Cited alongside, same era.
Compositional distributional semantics with long short term memory, 2015
Phong Le and Willem Zuidema · 2015
Cited alongside, same era.
Improving deep neural networks using softplus units
Hao Zheng, Zhanlei Yang, Wenju Liu, Jizhong Liang, and Yanpeng Li · 2015
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dandelion Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2015
Cited alongside, same era.
Rethinking the inception architecture for computer vision, 2015
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna · 2015
Improving deep learning by inverse square root linear units (isrlus), 2017
Brad Carlile, Guy Delamarter, Paul Kinney, Akiko Marti, and Brian Whitney · 2017
Later among the works it cites.
Searching for activation functions, 2017
Prajit Ramachandran, Barret Zoph, and Quoc V. Le · 2017
Later among the works it cites.
Deep voice 3: Scaling text-to-speech with convolutional sequence learning, 2017
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O. Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller · 2017
Later among the works it cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications, 2017
Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Later among the works it cites.
Activation functions: Comparison of trends in practice and research for deep learning, 2018
Chigozie Nwankpa, Winifred Ijomah, Anthony Gachagan, and Stephen Marshall · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Language modeling with gated convolutional networks, 2016
Yann N. Dauphin, Angela Fan, Michael Auli, and David Grangier · 2016
Cited alongside, same era.
Densely connected convolutional networks, 2016
Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q. Weinberger · 2016
Cited alongside, same era.
Lets keep it simple, using simple architectures to outperform deeper and more complex architectures, 2016
Seyyed Hossein Hasanpour, Mohammad Rouhani, Mohsen Fayyaz, and Mohammad Sabokrou · 2016
Cited alongside, same era.
Wide residual networks, 2016
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
E-swish: Adjusting activations to different network depths, 2018
Eric Alcaide · 2018
Later among the works it cites.
The quest for the golden activation function, 2018
Mina Basirat and Peter M. Roth · 2018
Later among the works it cites.
Mish: A self regularized non-monotonic activation function, 2019
Diganta Misra · 2019
Later among the works it cites.
Soft-root-sign activation function, 2020
Yuan Zhou, Dandan Li, Shuwei Huo, and Sun-Yuan Kung · 2020
Closest in time.
Tanhexp: A smooth activation function with high convergence speed for lightweight neural networks, 2020
Xinyu Liu and Xiaoguang Di · 2020
Closest in time.