Fetching the paper…
Reading the bibliography…
There are two widely known issues with properly training Recurrent Neural Networks, the vanishing and the exploding gradient problems detailed in Bengio et al.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J. (1986) · 1986
Earlier work this paper cites.
Generalization of backpropagation with application to a recurrent gas market model
Werbos, P. J. (1988) · 1988
Earlier work this paper cites.
Finding structure in time
Elman, J. (1990) · 1990
Earlier work this paper cites.
Adaptive synchronization of neural and physical oscillators
Doya, K. and Yoshizawa, S. (1991) · 1991
Earlier work this paper cites.
The problem of learning long-term dependencies in recurrent networks
Bengio, Y., Frasconi, P., and Simard, P. (1993) · 1993
Earlier work this paper cites.
Bifurcations of recurrent neural networks in gradient descent learning
Doya, K. (1993) · 1993
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Bengio, Y., Simard, P., and Frasconi, P. (1994) · 1994
Earlier work this paper cites.
Nonlinear Dynamics And Chaos: With Applications To Physics, Biology, Chemistry, And Engineering (Studies in Nonlinearity)
Strogatz, S. (1994) · 1994
Earlier work this paper cites.
Neural networks with adaptive learning rate and momentum terms
Moreira, M. and Fiesler, E. (1995) · 1995
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Cited alongside, same era.
Optimization and applications of echo state networks with leaky- integrator neurons
Jaeger, H., Lukosevicius, M., Popovici, D., and Siewert, U. (2007) · 2007
Cited alongside, same era.
A Novel Connectionist System for Unconstrained Handwriting Recognition
Graves, A., Liwicki, M., Fernandez, S., Bertolami, R., Bunke, H., and Schmidhuber, J. (2009) · 2009
Cited alongside, same era.
Reservoir computing approaches to recurrent neural network training
Lukoševičius, M. and Jaeger, H. (2009) · 2009
Cited alongside, same era.
Theano: a CPU and GPU math expression compiler
Bergstra, J., Breuleux, O., Bastien, F., Lamblin, P., Pascanu, R., Desjardins, G., Turian, J., Warde-Farley, D., and Bengio, Y. (2010) · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
A neurodynamical model for working memory
Pascanu, R. and Jaeger, H. (2011) · 2011
Later among the works it cites.
Generating text with recurrent neural networks
Sutskever, I., Martens, J., and Hinton, G. (2011) · 2011
Later among the works it cites.
Theano: new features and speed improvements
Bastien, F., Lamblin, P., Pascanu, R., Bergstra, J., Goodfellow, I., Bergeron, A., Bouchard, N., and Bengio, Y. (2012) · 2012
Closest in time.
Advances in optimizing recurrent networks
Bengio, Y., Boulanger-Lewandowski, N., and Pascanu, R. (2012) · 2012
Closest in time.
Modeling temporal dependencies in high-dimensional sequences: Application to polyphonic music generation and transcription
Boulanger-Lewandowski, N., Bengio, Y., and Vincent, P. (2012) · 2012
Closest in time.
Long short-term memory in echo state networks: Details of a simulation study
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Duchi, J. C., Hazan, E., and Singer, Y. (2011) · 2011
Cited alongside, same era.
Learning recurrent neural networks with Hessian-free optimization
Martens, J. and Sutskever, I. (2011) · 2011
Cited alongside, same era.
Empirical evaluation and combination of advanced language modeling techniques
Mikolov, T., Deoras, A., Kombrink, S., Burget, L., and Cernocky, J. (2011) · 2011
Cited alongside, same era.
Jaeger, H. (2012) · 2012
Closest in time.
Statistical Language Models based on Neural Networks
Mikolov, T. (2012) · 2012
Closest in time.
Subword language modeling with neural networks
Mikolov, T., Sutskever, I., Deoras, A., Le, H.-S., Kombrink, S., and Cernocky, J. (2012) · 2012
Closest in time.