Fetching the paper…
Reading the bibliography…
We stabilize the activations of Recurrent Neural Networks (RNNs) by penalizing the squared distance between successive hidden states' norms.
An analysis of noise in recurrent neural networks: convergence and generalization
Jim, Kam-Chuen, Giles, C. Lee, and Horne, Bill G · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Recurrent nets that time and count
Gers, Felix A. and Schmidhuber, Jürgen · 2000
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
Graves, Alex, Fernández, Santiago, Gomez, Faustino, and Schmidhuber, Jürgen · 2006
Earlier work this paper cites.
Theano: new features and speed improvements
Bastien, Frédéric, Lamblin, Pascal, Pascanu, Razvan, Bergstra, James, Goodfellow, Ian J., Bergeron, Arnaud, Bouchard, Nicolas, Warde-Farley, David, and Bengio, Yoshua · 2012
Earlier work this paper cites.
The human knowledge compression contest
Hutter, Marcus · 2012
Earlier work this paper cites.
Understanding the exploding gradient problem
Pascanu, Razvan, Mikolov, Tomas, and Bengio, Yoshua · 2012
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Graves, A., Mohamed, A.-R., and Hinton, G · 2013
Earlier work this paper cites.
Regularization and nonlinearities for neural language models: when are they needed?
Pachitariu, M. and Sahani, M · 2013
Cited alongside, same era.
Dropout improves recurrent neural networks for handwriting recognition
Pham, Vu, Kermorvant, Christopher, and Louradour, Jérôme · 2013
Cited alongside, same era.
Deep Speech: Scaling up end-to-end speech recognition
Hannun, A., Case, C., Casper, J., Catanzaro, B., Diamos, G., Elsen, E., Prenger, R., Satheesh, S., Sengupta, S., Coates, A., and Ng, A. Y · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik P. and Ba, Jimmy · 2014
Cited alongside, same era.
Zero-bias autoencoders and the benefits of co-adapting features
Konda, K., Memisevic, R., and Krueger, D · 2014
Cited alongside, same era.
Greff, Klaus, Srivastava, Rupesh Kumar, Koutník, Jan, Steunebrink, Bas R., and Schmidhuber, Jürgen · 2015
Closest in time.
Learning state representations with robotic priors
Jonschkowski, Rico and Brock, Oliver · 2015
Closest in time.
A simple way to initialize recurrent networks of rectified linear units
Le, Quoc V., Jaitly, Navdeep, and Hinton, Geoffrey E · 2015
Closest in time.
Learning acoustic frame labeling for speech recognition with recurrent neural networks
Sak, Hasim, Senior, Andrew, Rao, Kanishka, Irsoy, Ozan, Graves, Alex, Beaufays, Françoise, and Schalkwyk, Johan · 2015
Closest in time.
Blocks and fuel: Frameworks for deep learning
van Merriënboer, Bart, Bahdanau, Dzmitry, Dumoulin, Vincent, Serdyuk, Dmitriy, Warde-Farley, David, Chorowski, Jan, and Bengio, Yoshua · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Combining time- and frequency-domain convolution in convolutional neural network-based phone recognition
Tóth, László · 2014
Cited alongside, same era.
Recurrent neural network regularization
Zaremba, Wojciech, Sutskever, Ilya, and Vinyals, Oriol · 2014
Cited alongside, same era.
Closest in time.
Semantically conditioned lstm-based natural language generation for spoken dialogue systems
Wen, Tsung-Hsien, Gasic, Milica, Mrksic, Nikola, Su, Pei-hao, Vandyke, David, and Young, Steve J · 2015
Closest in time.
Building a large annotated corpus of english: The penn treebank
Marcus, Mitchell P., Marcinkiewicz, Mary Ann, and Santorini, Beatrice · 2017
Closest in time.