Fetching the paper…
Reading the bibliography…
Training recurrent neural networks (RNNs) is a hard problem due to degeneracies in the optimization landscape, a problem also known as vanishing/exploding gradients.
Bounds for iterates, inverses, spectral variation and fields of values of non-normal matrices
Peter Henrici · 1962
Earlier work this paper cites.
Untersuchungen zu dynamischen neuronalen Netzen
S. Hochreiter · 1991
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P. Frasconi · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Memory traces in dynamical systems
S. Ganguli, D. Huh, and H. Sompolinsky · 2008
Earlier work this paper cites.
Learning phrase representations using rnn encoder–decoder for statistical machine translation
K. Cho, B. van Merriënboer, Ç Gülçehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
A simple way to initialize recurrent networks of rectified linear units
Q.V. Le, N. Jaitly, and G.E. Hinton · 2015
Cited alongside, same era.
Unitary evolution recurrent neural networks
M. Arjovsky, A. Shah, and Y. Bengio · 2016
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (elus)
D.-A. Clevert, T. Unterthiner, and S. Hochreiter · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Recurrent orthogonal networks and long-memory tasks
Mikael Henaff, Arthur Szlam, and Yann LeCun · 2016
Cited alongside, same era.
Recurrent network models of sequence generation and memory
Kanaka Rajan, Christopher D Harvey, and David W Tank · 2016
Full-capacity unitary recurrent neural networks
S. Wisdom, T. Powers, J.R. Hershey, J. Le Roux, and L. Atlas · 2016
Later among the works it cites.
Dilated recurrent neural networks
S. Chang, Y. Zhang, W. Han, M. Yu, X. Guo, W. Tan, X. Cui, M. Witbrock, M.A. Hasegawa-Johnson, and T.S. Huang · 2017
Later among the works it cites.
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
Jeffrey Pennington, Samuel Schoenholz, and Surya Ganguli · 2017
Later among the works it cites.
An empirical evaluation of generic convolutional and recurrent networks for sequence modeling
Shaojie Bai, J. Zico Kolter, and Vladlen Koltun · 2018
Later among the works it cites.
Dynamical isometry and a mean field theory of rnns: Gating enables signal propagation in recurrent neural networks
Minmin Chen, Jeffrey Pennington, and Samuel Schoenholz · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
An analysis of neural language modeling at multiple scales
Stephen Merity, Nitish Shirish Keskar, and Richard Socher
Cited in the paper.
Regularizing and optimizing lstm language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher
Cited in the paper.
Giancarlo Kerg, Kyle Goyette, Maximilian Puelma Touzel, Gauthier Gidel, Eugene Vorontsov, Yoshua Bengio, and Guillaume Lajoie · 2019
Closest in time.