Fetching the paper…
Reading the bibliography…
Common nonlinear activation functions used in neural networks can cause training difficulties due to the saturation behavior of the activation function, which may hide dependencies that are not visible to vanilla-SGD (using first order gradients only).
Numerical Continuation Methods. An Introduction
Allgower, E. L. and Georg, K · 1980
Earlier work this paper cites.
Optimization by simulated annealing
Kirkpatrick, S., Jr., C. D. Gelatt, , and Vecchi, M. P · 1983
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Curriculum learning
Bengio, Yoshua, Louradour, Jerome, Collobert, Ronan, and Weston, Jason · 2009
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Nair, Vinod and Hinton, Geoffrey E · 2010
Earlier work this paper cites.
Deep sparse rectifier neural networks
Glorot, Xavier, Bordes, Antoine, and Bengio, Yoshua · 2011
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons
Bengio, Yoshua · 2013
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Yoshua, Léonard, Nicholas, and Courville, Aaron · 2013
Earlier work this paper cites.
Goodfellow, Ian J, Warde-Farley, David, Mirza, Mehdi, Courville, Aaron, and Bengio, Yoshua · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua · 2014
Cited alongside, same era.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Cho, Kyunghyun, Van Merriënboer, Bart, Gulcehre, Caglar, Bahdanau, Dzmitry, Bougares, Fethi, Schwenk, Holger, and Bengio, Yoshua · 2014
Cited alongside, same era.
Graves, Alex, Wayne, Greg, and Danihelka, Ivo · 2014
Cited alongside, same era.
Deep speech: Scaling up end-to-end speech recognition
Hannun, Awni, Case, Carl, Casper, Jared, Catanzaro, Bryan, Diamos, Greg, Elsen, Erich, Prenger, Ryan, Satheesh, Sanjeev, Sengupta, Shubho, Coates, Adam, et al · 2014
Cited alongside, same era.
Teaching machines to read and comprehend
Hermann, Karl Moritz, Kocisky, Tomas, Grefenstette, Edward, Espeholt, Lasse, Kay, Will, Suleyman, Mustafa, and Blunsom, Phil · 2015
Later among the works it cites.
Visualizing and understanding recurrent networks
Karpathy, Andrej, Johnson, Justin, and Fei-Fei, Li · 2015
Later among the works it cites.
A simple way to initialize recurrent networks of rectified linear units
Le, Quoc V, Jaitly, Navdeep, and Hinton, Geoffrey E · 2015
Later among the works it cites.
Deep learning
LeCun, Yann, Bengio, Yoshua, and Hinton, Geoffrey · 2015
Later among the works it cites.
Adding gradient noise improves learning for very deep networks
Neelakantan, Arvind, Vilnis, Luke, Le, Quoc V, Sutskever, Ilya, Kaiser, Lukasz, Kurach, Karol, and Martens, James · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Weston, Jason, Chopra, Sumit, and Bordes, Antoine · 2014
Cited alongside, same era.
Zaremba, Wojciech and Sutskever, Ilya · 2014
Cited alongside, same era.
Recurrent neural network regularization
Zaremba, Wojciech, Sutskever, Ilya, and Vinyals, Oriol · 2014
Cited alongside, same era.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Ge, Rong, Huang, Furong, Jin, Chi, and Yuan, Yang · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
Xu, Kelvin, Ba, Jimmy, Kiros, Ryan, Courville, Aaron, Salakhutdinov, Ruslan, Zemel, Richard, and Bengio, Yoshua · 2015
Later among the works it cites.
Describing videos by exploiting temporal structure
Yao, Li, Torabi, Atousa, Cho, Kyunghyun, Ballas, Nicolas, Pal, Christopher, Larochelle, Hugo, and Courville, Aaron · 2015
Later among the works it cites.
Training recurrent neural networks by diffusion
Mobahi, Hossein · 2016
Closest in time.