Fetching the paper…
Reading the bibliography…
Recurrent neural networks are powerful models for processing sequential data, but they are generally plagued by vanishing and exploding gradient problems.
The measure of the critical values of differentiable maps
A. Sard · 1942
Earlier work this paper cites.
Unitary triangularization of a nonsymmetric matrix
A. S. Householder · 1958
Earlier work this paper cites.
DARPA TIMIT acoustic-phonetic continous speech corpus
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, and D. S. Pallett · 1993
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P. Frasconi · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Heterogeneous acoustic measurements and multiple classifiers for speech recognition
A. K. Halberstadt · 1998
Earlier work this paper cites.
Perceptual evaluation of speech quality (PESQ): An objective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs, 2000
ITU-T P.862 · 2000
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies
S. Hochreiter, Y. Bengio, P. Frasconi, and J. Schmidhuber · 2001
Earlier work this paper cites.
Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs
A. Rix, J. Beerends, M. Hollier, and A. Hekstra · 2001
Cited alongside, same era.
VOICEBOX: Speech processing toolbox for MATLAB, 2002
M. Brookes · 2002
Cited alongside, same era.
Speech Enhancement: Theory and Practice
P. C. Loizou · 2007
Cited alongside, same era.
Lie groups, physics, and geometry: an introduction for physicists, engineers and chemists
R. Gilmore · 2008
Cited alongside, same era.
Notes on optimization on Stiefel manifolds
H. D. Tagare · 2011
Cited alongside, same era.
An algorithm for intelligibility prediction of time-frequency weighted noisy speech
C. Taal, R. Hendriks, R. Heusdens, and J. Jensen · 2011
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
A. M. Saxe, J. L. McClelland, and S. Ganguli · 2013
Later among the works it cites.
On the properties of neural machine translation: Encoder-decoder approaches
K. Cho, B. van Merriënboer, D. Bahdanau, and Y. Bengio · 2014
Later among the works it cites.
Recurrent models of visual attention
V. Mnih, N. Heess, A. Graves, and K. Kavukcuoglu · 2014
Later among the works it cites.
A simple way to initialize recurrent networks of rectified linear units
Q. V. Le, N. Jaitly, and G. E. Hinton · 2015
Later among the works it cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the difficulty of training Recurrent Neural Networks
R. Pascanu, T. Mikolov, and Y. Bengio · 2012
Cited alongside, same era.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude, 2012
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
Later among the works it cites.
Unitary Evolution Recurrent Neural Networks
M. Arjovsky, A. Shah, and Y. Bengio · 2016
Closest in time.
Theano: A Python framework for fast computation of mathematical expressions
Theano Development Team · 2016
Closest in time.