Fetching the paper…
Reading the bibliography…
Plain recurrent networks greatly suffer from the vanishing gradient problem while Gated Neural Networks (GNNs) such as Long-short Term Memory (LSTM) and Gated Recurrent Unit (GRU) deliver promising results in many sequence learning tasks through sophisticated network designs.
Learning complex, extended sequences using the principle of history compression
Jürgen Schmidhuber · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann Lecun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Geoffrey E. Hinton, Simon Osindero, and Yee-Whye Teh · 2006
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol · 2008
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Learning recurrent neural networks with hessian-free optimization
James Martens and Ilya Sutskever · 2011
Earlier work this paper cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
How to construct deep recurrent neural networks
Razvan Pascanu, Çaglar Gülçehre, Kyunghyun Cho, and Yoshua Bengio · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M. Saxe, James L. McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Cited alongside, same era.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Çaglar Gülçehre, KyungHyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
A simple way to initialize recurrent networks of rectified linear units
Tunable efficient unitary neural networks (EUNN) and their application to RNN
Li Jing, Yichen Shen, Tena Dubcek, John Peurifoy, Scott A. Skirlo, Max Tegmark, and Marin Soljacic · 2016
Later among the works it cites.
Phased lstm: Accelerating recurrent network training for long or event-based sequences
Daniel Neil, Michael Pfeiffer, and Shih-Chii Liu · 2016
Later among the works it cites.
Residual networks are exponential ensembles of relatively shallow networks
Andreas Veit, Michael J. Wilber, and Serge J. Belongie · 2016
Later among the works it cites.
Capacity and trainability in recurrent neural networks
Jasmine Collins, Jascha Sohl-Dickstein, and David Sussillo · 2017
Later among the works it cites.
Residual Connections Encourage Iterative Inference
S. Jastrzebski, D. Arpit, N. Ballas, V. Verma, T. Che, and Y. Bengio · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Quoc V. Le, Navdeep Jaitly, and Geoffrey E. Hinton · 2015
Cited alongside, same era.
Dmytro Mishkin and Jiri Matas · 2015
Cited alongside, same era.
Training very deep networks
Rupesh K Srivastava, Klaus Greff, and Juergen Schmidhuber · 2015
Cited alongside, same era.
Towards ai-complete question answering: A set of prerequisite toy tasks
Jason Weston, Antoine Bordes, Sumit Chopra, and Tomas Mikolov · 2015
Cited alongside, same era.
Highway and residual networks learn unrolled iterative estimation
Klaus Greff, Rupesh Kumar Srivastava, and Jürgen Schmidhuber · 2016
Cited alongside, same era.
Identity Mappings in Deep Residual Networks , pp. 630–645
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Later among the works it cites.
Gated Orthogonal Recurrent Units: On Learning to Forget
Li Jing, Cagla Gulcehre, John Peurifoy, Yichen Shen, Max Tegmark, Marin Soljačić, and Yoshua Bengio · 2017
Later among the works it cites.
Learning simpler language models with the differential state framework
Alexander G. Ororbia II, Tomas Mikolov, and David Reitter · 2017
Later among the works it cites.
Di Xie, Jiang Xiong, and Shiliang Pu · 2017
Later among the works it cites.
DiracNets: Training Very Deep Neural Networks Without Skip-Connections
Sergey Zagoruyko and Nikos Komodakis · 2017
Later among the works it cites.