Fetching the paper…
Reading the bibliography…
In this paper, we explore different ways to extend a recurrent neural network (RNN) to a \textit{deep} RNN.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J. (1986) · 1986
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Hornik, K., Stinchcombe, M., and White, H. (1989) · 1989
Earlier work this paper cites.
Learning complex, extended sequences using the principle of history compression
Schmidhuber, J. (1992) · 1992
Earlier work this paper cites.
Building a large annotated corpus of english: The Penn Treebank
Marcus, M. P., Marcinkiewicz, M. A., and Santorini, B. (1993) · 1993
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Bengio, Y., Simard, P., and Frasconi, P. (1994) · 1994
Earlier work this paper cites.
Hierarchical recurrent neural networks for long-term dependencies
El Hihi, S. and Bengio, Y. (1996) · 1996
Earlier work this paper cites.
An unsupervised ensemble learning method for nonlinear dynamic state-space models
Valpola, H. and Karhunen, J. (2002) · 2002
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Hinton, G. E. and Salakhutdinov, R. (2006) · 2006
Earlier work this paper cites.
Discovering multiscale dynamical features with hierarchical echo state networks
Jaeger, H. (2007) · 2007
Earlier work this paper cites.
Learning deep architectures for AI
Bengio, Y. (2009) · 2009
Earlier work this paper cites.
Measuring invariances in deep networks
Goodfellow, I., Le, Q., Saxe, A., and Ng, A. (2009) · 2009
Earlier work this paper cites.
A novel connectionist system for improved unconstrained handwriting recognition
Graves, A., Liwicki, M., Fernandez, S., Bertolami, R., Bunke, H., and Schmidhuber, J. (2009) · 2009
Earlier work this paper cites.
Gp-bayesfilters: Bayesian filtering using gaussian process prediction and observation models
Ko, J. and Dieter, F. (2009) · 2009
Earlier work this paper cites.
Theano: a CPU and GPU math expression compiler
Bergstra, J., Breuleux, O., Bastien, F., Lamblin, P., Pascanu, R., Desjardins, G., Turian, J., Warde-Farley, D., and Bengio, Y. (2010) · 2010
Earlier work this paper cites.
Deep belief networks are compact universal approximators
Le Roux, N. and Bengio, Y. (2010) · 2010
Cited alongside, same era.
Deep learning via Hessian-free optimization
Martens, J. (2010) · 2010
Cited alongside, same era.
Recurrent neural network based language model
Mikolov, T., Karafiát, M., Burget, L., Cernocky, J., and Khudanpur, S. (2010) · 2010
Cited alongside, same era.
Shallow vs. deep sum-product networks
Delalleau, O. and Bengio, Y. (2011) · 2011
Cited alongside, same era.
Domain adaptation for large-scale sentiment classification: A deep learning approach
Glorot, X., Bordes, A., and Bengio, Y. (2011b) · 2011
Cited alongside, same era.
Practical variational inference for neural networks
Graves, A. (2011) · 2011
Cited alongside, same era.
Deep learning made easier by linear transformations in perceptrons
Raiko, T., Valpola, H., and LeCun, Y. (2012) · 2012
Later among the works it cites.
On fast dropout and its applicability to recurrent networks
Bayer, J., Osendorfer, C., Korhammer, D., Chen, N., Urban, S., and van der Smagt, P. (2013) · 2013
Closest in time.
Better mixing via deep representations
Bengio, Y., Mesnil, G., Dauphin, Y., and Rifai, S. (2013) · 2013
Closest in time.
A new method for learning deep recurrent neural networks
Chen, J. and Deng, L. (2013) · 2013
Closest in time.
Maxout networks
Goodfellow, I. J., Warde-Farley, D., Mirza, M., Courville, A., and Bengio, Y. (2013) · 2013
Closest in time.
Generating sequences with recurrent neural networks
Graves, A. (2013) · 2013
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The Neural Autoregressive Distribution Estimator
Larochelle, H. and Murray, I. (2011) · 2011
Cited alongside, same era.
Learning recurrent neural networks with Hessian-free optimization
Martens, J. and Sutskever, I. (2011) · 2011
Cited alongside, same era.
Extensions of recurrent neural network language model
Mikolov, T., Kombrink, S., Burget, L., Cernocky, J., and Khudanpur, S. (2011) · 2011
Cited alongside, same era.
Generating text with recurrent neural networks
Sutskever, I., Martens, J., and Hinton, G. (2011) · 2011
Cited alongside, same era.
Theano: new features and speed improvements
Bastien, F., Lamblin, P., Pascanu, R., Bergstra, J., Goodfellow, I. J., Bergeron, A., Bouchard, N., and Bengio, Y. (2012) · 2012
Cited alongside, same era.
Modeling temporal dependencies in high-dimensional sequences: Application to polyphonic music generation and transcription
Boulanger-Lewandowski, N., Bengio, Y., and Vincent, P. (2012) · 2012
Cited alongside, same era.
Speech recognition with deep recurrent neural networks
Graves, A., Mohamed, A., and Hinton, G. (2013) · 2013
Closest in time.
Learned-norm pooling for deep feedforward and recurrent neural networks
Gulcehre, C., Cho, K., Pascanu, R., and Bengio, Y. (2013) · 2013
Closest in time.
Training and analysing deep recurrent neural networks
Hermans, M. and Schrauwen, B. (2013) · 2013
Closest in time.
Revisiting natural gradient for deep networks
Pascanu, R. and Bengio, Y. (2013) · 2013
Closest in time.
On the difficulty of training recurrent neural networks
Pascanu, R., Mikolov, T., and Bengio, Y. (2013a) · 2013
Closest in time.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G. (2013) · 2013
Closest in time.
Recurrent convolutional neural networks for scene labeling
Pinheiro, P. and Collobert, R. (2014) · 2014
Closest in time.