Fetching the paper…
Reading the bibliography…
We propose a simple extension to the ReLU-family of activation functions that allows them to shift the mean activation across a layer towards zero.
Finding structure in time
Jeffrey L Elman · 1990
Earlier work this paper cites.
Eigenvalues of covariance matrices: Application to neural-network learning
Yann Le Cun, Ido Kanter, and Sara A Solla · 1991
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell P Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 1993
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Y. Simard, and Paolo Frasconi · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-ichi Amari · 1998
Earlier work this paper cites.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 1998
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E. Hinton · 2010
Earlier work this paper cites.
Deep sparse rectifier neural networks
Xavier Glorot, Antoine Bordes, and Yoshua Bengio · 2011
Earlier work this paper cites.
Large text compression benchmark: About the test data (http://mattmahoney.net/dc/textdata), 2011
Matt Mahoney · 2011
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Alex Graves · 2013
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Andrew L Maas, Awni Y Hannun, and Andrew Y Ng · 2013
Earlier work this paper cites.
Regularization and nonlinearities for neural language models: when are they needed?
Marius Pachitariu and Maneesh Sahani · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
Dropout improves recurrent neural networks for handwriting recognition
Vu Pham, Christopher Kermorvant, and Jérôme Louradour · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Recurrent neural network regularization
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals · 2014
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2015
Cited alongside, same era.
Deep networks with stochastic depth
Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q. Weinberger · 2016
Later among the works it cites.
Multiplicative LSTM for sequence modelling
Ben Krause, Liang Lu, Iain Murray, and Steve Renals · 2016
Later among the works it cites.
Zoneout: Regularizing rnns by randomly preserving hidden activations
David Krueger, Tegan Maharaj, János Kramár, Mohammad Pezeshki, Nicolas Ballas, Nan Rosemary Ke, Anirudh Goyal, Yoshua Bengio, Hugo Larochelle, Aaron C. Courville, and Chris Pal · 2016
Later among the works it cites.
Batch normalized recurrent neural networks
César Laurent, Gabriel Pereyra, Philemon Brakel, Ying Zhang, and Yoshua Bengio · 2016
Later among the works it cites.
Path-normalized optimization of recurrent neural networks with relu activations
Behnam Neyshabur, Yuhuai Wu, Ruslan R Salakhutdinov, and Nati Srebro · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Regularizing rnns by stabilizing activations
David Krueger and Roland Memisevic · 2015
Cited alongside, same era.
A simple way to initialize recurrent networks of rectified linear units
Quoc V. Le, Navdeep Jaitly, and Geoffrey E. Hinton · 2015
Cited alongside, same era.
Dmytro Mishkin and Jiri Matas · 2015
Cited alongside, same era.
Rnndrop: A novel dropout for rnns in asr
Taesup Moon, Heeyoul Choi, Hoshik Lee, and Inchul Song · 2015
Cited alongside, same era.
Empirical evaluation of rectified activations in convolutional network
Bing Xu, Naiyan Wang, Tianqi Chen, and Mu Li · 2015
Cited alongside, same era.
Lei Jimmy Ba, Ryan Kiros, and Geoffrey E. Hinton · 2016
Cited alongside, same era.
Later among the works it cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Diederik P Kingma · 2016
Later among the works it cites.
Recurrent dropout without memory loss
Stanislau Semeniuta, Aliaksei Severyn, and Erhardt Barth · 2016
Later among the works it cites.
Understanding and improving convolutional neural networks via concatenated rectified linear units
Wenling Shang, Kihyuk Sohn, Diogo Almeida, and Honglak Lee · 2016
Later among the works it cites.
On multiplicative integration with recurrent neural networks
Yuhuai Wu, Saizheng Zhang, Ying Zhang, Yoshua Bengio, and Ruslan Salakhutdinov · 2016
Later among the works it cites.
Sergey Zagoruyko and Nikos Komodakis · 2016
Later among the works it cites.
Julian G. Zilly, Rupesh Kumar Srivastava, Jan Koutník, and Jürgen Schmidhuber · 2016
Later among the works it cites.
The shattered gradients problem: If resnets are the answer, then what is the question?
David Balduzzi, Marcus Frean, Lennox Leary, J. P. Lewis, Kurt Wan-Duo Ma, and Brian McWilliams · 2017
Closest in time.
Learning simpler language models with the delta recurrent neural network framework
Alexander G. Ororbia II, Tomas Mikolov, and David Reitter · 2017
Closest in time.
Self-normalizing neural networks
Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter · 2017
Closest in time.
Yanzhao Zhou, Qixiang Ye, Qiang Qiu, and Jianbin Jiao · 2017
Closest in time.