Fetching the paper…
Reading the bibliography…
Learning long term dependencies in recurrent networks is difficult due to vanishing and exploding gradients.
Learning representations by back-propagating errors
D. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Learning to forget: Continual prediction with LSTM
F. A. Gers, J. Schmidhuber, and F. Cummins · 2000
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies
S. Hochreiter, Y. Bengio, P. Frasconi, and J. Schmidhuber · 2001
Earlier work this paper cites.
Learning precise timing with lstm recurrent networks
F. A. Gers, N. N. Schraudolph, and J. Schmidhuber · 2003
Earlier work this paper cites.
A novel connectionist system for unconstrained handwriting recognition
A. Graves, M. Liwicki, S. Fernández, R. Bertolami, H. Bunke, and J. Schmidhuber · 2009
Earlier work this paper cites.
Deep learning via Hessian-free optimization
J. Martens · 2010
Earlier work this paper cites.
Rectified Linear Units improve Restricted Boltzmann Machines
V. Nair and G. Hinton · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Learning recurrent neural networks with Hessian-Free optimization
J. Martens and I. Sutskever · 2011
Earlier work this paper cites.
The kaldi speech recognition toolkit
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely · 2011
Earlier work this paper cites.
Generating text with recurrent neural networks
I. Sutskever, J. Martens, and G. E. Hinton · 2011
Earlier work this paper cites.
Context-dependent pre-trained deep neural networks for large vocabulary speech recognition
G. E. Dahl, D. Yu, L. Deng, and A. Acero · 2012
Earlier work this paper cites.
Large scale distributed deep networks
J. Dean, G. S. Corrado, R. Monga, K. Chen, M. Devin, Q. V. Le, M. Z. Mao, M. A. Ranzato, A. Senior, P. Tucker, K. Yang, and A. Y. Ng · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
G. Hinton · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
Building high-level features using large scale unsupervised learning
Q. V. Le, M. A. Ranzato, R. Monga, M. Devin, K. Chen, G. S. Corrado, J. Dean, and A. Y. Ng · 2012
Cited alongside, same era.
Training deep and recurrent neural networks with Hessian-Free optimization
J. Martens and I. Sutskever · 2012
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Later among the works it cites.
On rectified linear units for speech processing
M. Zeiler, M. Ranzato, R. Monga, M. Mao, K. Yang, Q. V. Le, P. Nguyen, A. Senior, V. Vanhoucke, and J. Dean · 2013
Later among the works it cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Later among the works it cites.
Towards end-to-end speech recognition with recurrent neural networks
A. Graves and N. Jaitly · 2014
Later among the works it cites.
Exploring Deep Learning Methods for discovering features in speech signals
N. Jaitly · 2014
Later among the works it cites.
Unifying visual-semantic embeddings with multimodal neural language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Pascanu, T. Mikolov, and Y. Bengio · 2012
Cited alongside, same era.
One billion word benchmark for measuring progress in statistical language modeling
C. Chelba, T. Mikolov, M. Schuster, Q. Ge, T. Brants, and P. Koehn · 2013
Cited alongside, same era.
Generating sequences with recurrent neural networks
A. Graves · 2013
Cited alongside, same era.
Generating sequences with recurrent neural networks
A. Graves · 2013
Cited alongside, same era.
Hybrid speech recognition with deep bidirectional lstm
A. Graves, N. Jaitly, and A-R. Mohamed · 2013
Cited alongside, same era.
Speech recognition with deep recurrent neural networks
A. Graves, A-R. Mohamed, and G. Hinton · 2013
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
A. M. Saxe, J. L. McClelland, and S. Ganguli · 2013
Cited alongside, same era.
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2014
Later among the works it cites.
Addressing the rare word problem in neural machine translation
T. Luong, I. Sutskever, Q. V. Le, O. Vinyals, and W. Zaremba · 2014
Later among the works it cites.
Learning longer memory in recurrent neural networks
T. Mikolov, A. Joulin, S. Chopra, M. Mathieu, and M. A. Ranzato · 2014
Later among the works it cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Later among the works it cites.
O. Vinyals, L. Kaiser, T. Koo, S. Petrov, I. Sutskever, and G. Hinton · 2014
Later among the works it cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2014
Later among the works it cites.
W. Zaremba and I. Sutskever · 2014
Later among the works it cites.
Random walk intialization for training very deep networks
D. Sussillo and L. F. Abbott · 2015
Closest in time.