Fetching the paper…
Reading the bibliography…
Many sequential processing tasks require complex nonlinear transition functions from one step to the next.
Über die Abgrenzung der Eigenwerte einer Matrix
Geršgorin, S · 1931
Earlier work this paper cites.
The representation of the cumulative rounding error of an algorithm as a taylor expansion of the local rounding errors
Linnainmaa, S · 1970
Earlier work this paper cites.
Taylor expansion of the accumulated rounding error
Linnainmaa, Seppo · 1976
Earlier work this paper cites.
System Modeling and Optimization: Proceedings of the 10th IFIP Conference New York City, USA, August 31 – September 4, 1981 , chapter Applications of advances in nonlinear sensitivity analysis, pp. 762–770
Werbos, Paul J · 1982
Earlier work this paper cites.
The utility driven dynamic error propagation network
Robinson, A. J. and Fallside, F · 1987
Earlier work this paper cites.
Generalization of backpropagation with application to a recurrent gas market model
Werbos, Paul J · 1988
Earlier work this paper cites.
Complexity of exact gradient computation algorithms for recurrent neural networks
Williams, R. J · 1989
Earlier work this paper cites.
Untersuchungen zu dynamischen neuronalen Netzen
Hochreiter, S · 1991
Earlier work this paper cites.
Reinforcement learning in markovian and non-markovian environments
Schmidhuber, Jürgen · 1991
Earlier work this paper cites.
Learning complex, extended sequences using the principle of history compression
Schmidhuber, Jürgen · 1992
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Bengio, Yoshua, Simard, Patrice, and Frasconi, Paolo · 1994
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Learning to forget: Continual prediction with lstm
Gers, Felix A., Schmidhuber, Jürgen, and Cummins, Fred · 2000
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies
Hochreiter, S., Bengio, Y., Frasconi, P., and Schmidhuber, J · 2001
Earlier work this paper cites.
Scaling learning algorithms towards ai
Bengio, Yoshua and LeCun, Yann · 2007
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, Tomas, Karafiát, Martin, Burget, Lukas, Cernockỳ, Jan, and Khudanpur, Sanjeev · 2010
Earlier work this paper cites.
Modeling temporal dependencies in high-dimensional sequences: Application to polyphonic music generation and transcription
Boulanger-Lewandowski, N., Bengio, Y., and Vincent, P · 2012
Earlier work this paper cites.
The human knowledge compression contest
Hutter, M · 2012
Cited alongside, same era.
Context dependent recurrent neural network language model
Mikolov, Tomas and Zweig, Geoffrey · 2012
Cited alongside, same era.
Generating sequences with recurrent neural networks
Graves, A · 2013
Cited alongside, same era.
How to construct deep recurrent neural networks
Pascanu, R., Gulcehre, C., Cho, K., and Bengio, Y · 2013
Cited alongside, same era.
First experiments with powerplay
Srivastava, Rupesh Kumar, Steunebrink, Bas R, and Schmidhuber, Jürgen · 2013
Cited alongside, same era.
On the complexity of neural network classifiers: A comparison between shallow and deep architectures
Bianchini, Monica and Scarselli, Franco · 2014
Cited alongside, same era.
Hierarchical Multiscale Recurrent Neural Networks
Chung, J., Ahn, S., and Bengio, Y · 2016
Closest in time.
Recurrent Batch Normalization
Cooijmans, T., Ballas, N., Laurent, C., Gülçehre, Ç., and Courville, A · 2016
Closest in time.
Adaptive Computation Time for Recurrent Neural Networks
Graves, A · 2016
Closest in time.
Highway and residual networks learn unrolled iterative estimation
Greff, Klaus, Srivastava, Rupesh K, and Schmidhuber, Jürgen · 2016
Closest in time.
HyperNetworks
Ha, D., Dai, A., and Le, Q. V · 2016
Closest in time.
Tying Word Vectors and Word Classifiers: A Loss Framework for Language Modeling
Inan, H., Khosravi, K., and Socher, R · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cho, Kyunghyun, Van Merriënboer, Bart, Gulcehre, Caglar, Bahdanau, Dzmitry, Bougares, Fethi, Schwenk, Holger, and Bengio, Yoshua · 2014
Cited alongside, same era.
Recurrent Neural Network Regularization
Zaremba, W., Sutskever, I., and Vinyals, O · 2014
Cited alongside, same era.
Gated feedback recurrent neural networks
Chung, Junyoung, Gulcehre, Caglar, Cho, Kyunghyun, and Bengio, Yoshua · 2015
Cited alongside, same era.
A theoretically grounded application of dropout in recurrent neural networks
Gal, Yarin · 2015
Cited alongside, same era.
Greff, K., Srivastava, R. K., Koutník, J., Steunebrink, B. R, and Schmidhuber, J · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Cited alongside, same era.
Improved learning through augmenting the loss, 2016
Inan, Hakan and Khosravi, Khashayar · 2016
Closest in time.
Exploring the limits of language modeling
Jozefowicz, Rafal, Vinyals, Oriol, Schuster, Mike, Shazeer, Noam, and Wu, Yonghui · 2016
Closest in time.
Multiplicative LSTM for sequence modelling
Krause, B., Lu, L., Murray, I., and Renals, S · 2016
Closest in time.
Layer Normalization
Lei Ba, J., Kiros, J. R., and Hinton, G. E · 2016
Closest in time.
Pointer Sentinel Mixture Models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Closest in time.
Using the Output Embedding to Improve Language Models
Press, O. and Wolf, L · 2016
Closest in time.
On Multiplicative Integration with Recurrent Neural Networks
Wu, Y., Zhang, S., Zhang, Y., Bengio, Y., and Salakhutdinov, R · 2016
Closest in time.
Highway long short-term memory RNNS for distant speech recognition
Zhang, Yu, Chen, Guoguo, Yu, Dong, Yao, Kaisheng, Khudanpur, Sanjeev, and Glass, James · 2016
Closest in time.
Neural architecture search with reinforcement learning
Zoph, Barret and Le, Quoc V · 2016
Closest in time.
Building a large annotated corpus of english: The penn treebank
Marcus, Mitchell P., Marcinkiewicz, Mary Ann, and Santorini, Beatrice · 2017
Closest in time.