Fetching the paper…
Reading the bibliography…
Recurrent Neural Networks (RNNs) are rich models for the processing of sequential data.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J. (1986) · 1986
Earlier work this paper cites.
Untersuchungen zu dynamischen neuronalen netzen
Hochreiter, S. (1991) · 1991
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Bengio, Y., Simard, P., and Frasconi, P. (1994) · 1994
Earlier work this paper cites.
Neural networks for pattern recognition
Bishop, C. M. (1995) · 1995
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
New extension of the kalman filter to nonlinear systems
Julier, S. J. and Uhlmann, J. K. (1997) · 1997
Earlier work this paper cites.
Elements of large-sample theory
Lehmann, E. L. (1999) · 1999
Earlier work this paper cites.
On the approximation capability of recurrent neural networks
Hammer, B. (2000) · 2000
Earlier work this paper cites.
Adaptive nonlinear system identification with echo state networks
Jäger, H. et al. (2003) · 2003
Earlier work this paper cites.
Pattern recognition and machine learning
Bishop, C. M. et al. (2006) · 2006
Earlier work this paper cites.
Unconstrained online handwriting recognition with recurrent neural networks
Graves, A., Fernández, S., Liwicki, M., Bunke, H., and Schmidhuber, J. (2008) · 2008
Earlier work this paper cites.
Evaluation of multiple-f0 estimation and tracking systems
Bay, M., Ehmann, A. F., and Downie, J. S. (2009) · 2009
Earlier work this paper cites.
Theano: a CPU and GPU math expression compiler
Bergstra, J., Breuleux, O., Bastien, F., Lamblin, P., Pascanu, R., Desjardins, G., Turian, J., Warde-Farley, D., and Bengio, Y. (2010) · 2010
Cited alongside, same era.
Recurrent neural network based language model
Mikolov, T., Karafiát, M., Burget, L., Cernocky, J., and Khudanpur, S. (2010) · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y. (2011) · 2011
Cited alongside, same era.
Practical variational inference for neural networks
Graves, A. (2011) · 2011
Cited alongside, same era.
Learning recurrent neural networks with hessian-free optimization
Martens, J. and Sutskever, I. (2011) · 2011
Cited alongside, same era.
The manifold tangent classifier
Rifai, S., Dauphin, Y. N., Vincent, P., Bengio, Y., and Muller, X. (2011) · 2011
Cited alongside, same era.
Lecture 6.5 - rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G. (2012) · 2012
Later among the works it cites.
Training neural networks with implicit variance
Bayer, J., Osendorfer, C., Urban, S., et al. (2013) · 2013
Closest in time.
High-dimensional sequence transduction
Boulanger-Lewandowski, N., Bengio, Y., and Vincent, P. (2013) · 2013
Closest in time.
Generating sequences with recurrent neural networks
Graves, A. (2013) · 2013
Closest in time.
Speech recognition with deep recurrent neural networks
Graves, A., Mohamed, A.-r., and Hinton, G. (2013) · 2013
Closest in time.
Regularization and nonlinearities for neural language models: when are they needed?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generating text with recurrent neural networks
Sutskever, I., Martens, J., and Hinton, G. (2011) · 2011
Cited alongside, same era.
Advances in optimizing recurrent networks
Bengio, Y., Boulanger-Lewandowski, N., and Pascanu, R. (2012) · 2012
Cited alongside, same era.
Random search for hyper-parameter optimization
Bergstra, J. and Bengio, Y. (2012) · 2012
Cited alongside, same era.
Improving neural networks by preventing co-adaptation of feature detectors
Hinton, G. E., Srivastava, N., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. R. (2012) · 2012
Cited alongside, same era.
Matrix analysis
Horn, R. A. and Johnson, C. R. (2012) · 2012
Cited alongside, same era.
On the difficulty of training recurrent neural networks
Pascanu, R., Mikolov, T., and Bengio, Y. (2012) · 2012
Cited alongside, same era.
Pachitariu, M. and Sahani, M. (2013) · 2013
Closest in time.
How to construct deep recurrent neural networks
Pascanu, R., Gulcehre, C., Cho, K., and Bengio, Y. (2013) · 2013
Closest in time.
Training Recurrent Neural Networks
Sutskever, I. (2013) · 2013
Closest in time.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G. (2013) · 2013
Closest in time.
Dropout training as adaptive regularization
Wager, S., Wang, S., and Liang, P. (2013) · 2013
Closest in time.
Fast dropout training
Wang, S. and Manning, C. (2013) · 2013
Closest in time.
On rectified linear units for speech processing
Zeiler, M., Ranzato, M., Monga, R., Mao, M., Yang, K., Le, Q., Nguyen, P., Senior, A., Vanhoucke, V., Dean, J., et al. (2013) · 2013
Closest in time.