Fetching the paper…
Reading the bibliography…
The choice of activation functions in deep networks has a significant effect on the training dynamics and task performance.
Evolutionary principles in self-referential learning
Jurgen Schmidhuber · 1987
Earlier work this paper cites.
Meta-neural networks that learn by learning
Devang K Naik and RJ Mammone · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Digital selection and analogue amplification coexist in a cortex-inspired silicon circuit
Richard HR Hahnloser, Rahul Sarpeshkar, Misha A Mahowald, Rodney J Douglas, and H Sebastian Seung · 2000
Earlier work this paper cites.
What is the best multi-stage architecture for object recognition?
Kevin Jarrett, Koray Kavukcuoglu, Yann LeCun, et al · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E Hinton · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Learning to learn
Sebastian Thrun and Lorien Pratt · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Maxout networks
Ian J Goodfellow, David Warde-Farley, Mehdi Mirza, Aaron Courville, and Yoshua Bengio · 2013
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Andrew L Maas, Awni Y Hannun, and Andrew Y Ng · 2013
Earlier work this paper cites.
Learning activation functions to improve deep neural networks
Forest Agostinelli, Matthew Hoffman, Peter Sadowski, and Pierre Baldi · 2014
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Cited alongside, same era.
Rupesh Kumar Srivastava, Klaus Greff, and Jürgen Schmidhuber · 2015
Cited alongside, same era.
Empirical evaluation of rectified activations in convolutional network
Bing Xu, Naiyan Wang, Tianqi Chen, and Mu Li · 2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Cited alongside, same era.
Layer normalization
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Cited alongside, same era.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le · 2016
Later among the works it cites.
Neural optimizer search with reinforcement learning
Irwan Bello, Barret Zoph, Vijay Vasudevan, and Quoc V Le · 2017
Closest in time.
Reinforcement learning for architecture search by network transformation
Han Cai, Tianyao Chen, Weinan Zhang, Yong Yu, and Jun Wang · 2017
Closest in time.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
Stefan Elfwing, Eiji Uchibe, and Kenji Doya · 2017
Closest in time.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier · 2016
Cited alongside, same era.
Rl2: Fast reinforcement learning via slow reinforcement learning
Yan Duan, John Schulman, Xi Chen, Peter L Bartlett, Ilya Sutskever, and Pieter Abbeel · 2016
Cited alongside, same era.
David Ha, Andrew Dai, and Quoc V Le · 2016
Cited alongside, same era.
Bridging nonlinearities and stochastic regularizers with gaussian error linear units
Dan Hendrycks and Kevin Gimpel · 2016
Cited alongside, same era.
Taming the waves: sine as activation function in deep neural networks
Giambattista Parascandolo, Heikki Huttunen, and Tuomas Virtanen · 2016
Cited alongside, same era.
Optimization as a model for few-shot learning
Sachin Ravi and Hugo Larochelle · 2016
Cited alongside, same era.
Understanding and improving convolutional neural networks via concatenated rectified linear units
Wenling Shang, Kihyuk Sohn, Diogo Almeida, and Honglak Lee · 2016
Cited alongside, same era.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Closest in time.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Kilian Q Weinberger, and Laurens van der Maaten · 2017
Closest in time.
Self-normalizing neural networks
Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter · 2017
Closest in time.
Learnable pooling with context gating for video classification
Antoine Miech, Ivan Laptev, and Josef Sivic · 2017
Closest in time.
Flexible rectified linear units for improving convolutional neural networks
Suo Qiu and Bolun Cai · 2017
Closest in time.
Large-scale evolution of image classifiers
Esteban Real, Sherry Moore, Andrew Selle, Saurabh Saxena, Yutaka Leon Suematsu, Quoc Le, and Alex Kurakin · 2017
Closest in time.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Closest in time.
Inception-v4, inception-resnet and the impact of residual connections on learning
Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi · 2017
Closest in time.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Closest in time.
Practical network blocks design with q-learning
Zhao Zhong, Junjie Yan, and Cheng-Lin Liu · 2017
Closest in time.
Deep interest network for click-through rate prediction
Guorui Zhou, Chengru Song, Xiaoqiang Zhu, Xiao Ma, Yanghui Yan, Xingya Dai, Han Zhu, Junqi Jin, Han Li, and Kun Gai · 2017
Closest in time.
Learning transferable architectures for scalable image recognition
Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V Le · 2017
Closest in time.