Fetching the paper…
Reading the bibliography…
The vanishing and exploding gradient problems are well-studied obstacles that make it difficult for recurrent neural networks to learn long-term time dependencies.
Learning long-term dependencies with gradient descent is difficult
Bengio, Yoshua, Simard, Patrice, and Frasconi, Paolo · 1994
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Matrix computations , volume 3
Golub, Gene H and Van Loan, Charles F · 2012
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, Andrew M., McClelland, James L., and Ganguli, Surya · 2013
Cited alongside, same era.
Unitary evolution recurrent neural networks
Arjovsky, Martín, Shah, Amar, and Bengio, Yoshua · 2015
Cited alongside, same era.
A simple way to initialize recurrent networks of rectified linear units
Le, Quoc V., Jaitly, Navdeep, and Hinton, Geoffrey E · 2015
Cited alongside, same era.
Improving training of deep neural networks via singular value bounding
Jia, Kui · 2016
Closest in time.
Full-capacity unitary recurrent neural networks
Wisdom, Scott, Powers, Thomas, Hershey, John, Le Roux, Jonathan, and Atlas, Les · 2016
Closest in time.
Zilly, Julian Georg, Srivastava, Rupesh Kumar, Koutník, Jan, and Schmidhuber, Jürgen · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…