Fetching the paper…
Reading the bibliography…
Recurrent neural networks (RNNs) are particularly well-suited for modeling long-term dependencies in sequential data, but are notoriously hard to train because the error backpropagated in time either vanishes or explodes at an exponential rate.
A comprehensive introduction to differential geometry
M. Spivak · 1970
Earlier work this paper cites.
Nonlinear Systems Analysis
M. Vidyasagar · 1978
Earlier work this paper cites.
Inexact newton methods
R. S. Dembo, S. C. Eisenstat, and T. Steihaug · 1982
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P. Frasconi · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Numerical Methods for Ordinary Differential Equations
J. Butcher · 2003
Earlier work this paper cites.
Optimization and applications of echo state networks with leaky-integrator neurons
H. Jaeger, M. Lukosevicius, D. Popovici, and U. Siewert · 2007
Earlier work this paper cites.
Comparative study on classifying human activities with miniature inertial and magnetic sensors
K. Altun, B. Barshan, and O. Tunçel · 2010
Earlier work this paper cites.
Human activity recognition on smartphones using a multiclass hardware-friendly support vector machine
D. Anguita, A. Ghio, L. Oneto, X. Parra, and J. L. Reyes-Ortiz · 2012
Earlier work this paper cites.
Advances in optimizing recurrent networks
Y. Bengio, N. Boulanger-Lewandowski, and R. Pascanu · 2013
Earlier work this paper cites.
Training and analysing deep recurrent neural networks
M. Hermans and B. Schrauwen · 2013
Earlier work this paper cites.
Hidden factors and hidden topics: Understanding rating dimensions with review text
J. McAuley and J. Leskovec · 2013
Earlier work this paper cites.
How to construct deep recurrent neural networks
R. Pascanu, C. Gulcehre, K. Cho, and Y. Bengio · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
R. Pascanu, T. Mikolov, and Y. Bengio · 2013
Earlier work this paper cites.
How to construct deep recurrent neural networks
R. Pascanu, Çaglar Gülçehre, K. Cho, and Y. Bengio · 2013
Cited alongside, same era.
On the properties of neural machine translation: Encoder-decoder approaches
K. Cho, B. Van Merriënboer, D. Bahdanau, and Y. Bengio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Improving performance of recurrent neural network with relu nonlinearity
S. S. Talathi and A. Vartak · 2015
Cited alongside, same era.
Unitary evolution recurrent neural networks
M. Arjovsky, A. Shah, and Y. Bengio · 2016
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Later among the works it cites.
Yelp dataset challenge
I. Yelp · 2017
Later among the works it cites.
Recurrent highway networks
J. G. Zilly, R. K. Srivastava, J. Koutník, and J. Schmidhuber · 2017
Later among the works it cites.
Neural ordinary differential equations
T. Q. Chen, Y. Rubanova, J. Bettencourt, and D. K. Duvenaud · 2018
Later among the works it cites.
Frage: frequency-agnostic word representation
C. Gong, D. He, X. Tan, T. Qin, L. Wang, and T.-Y. Liu · 2018
Later among the works it cites.
Fastgrnn: A fast, accurate, stable and tiny kilobyte sized gated recurrent neural network
A. Kusupati, M. Singh, K. Bhatia, A. Kumar, P. Jain, and M. Varma · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Collins, J. Sohl-Dickstein, and D. Sussillo · 2016
Cited alongside, same era.
T. Cooijmans, N. Ballas, C. Laurent, Ç. Gülçehre, and A. Courville · 2016
Cited alongside, same era.
Efficient orthogonal parametrisation of recurrent neural networks using householder reflections
Z. Mhammedi, A. D. Hellicar, A. Rahman, and J. Bailey · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. W. Senior, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Skip rnn: Learning to skip state updates in recurrent neural networks
V. Campos, B. Jou, X. Giró-i Nieto, J. Torres, and S.-F. Chang · 2017
Cited alongside, same era.
Dilated recurrent neural networks
S. Chang, Y. Zhang, W. Han, M. Yu, X. Guo, W. Tan, X. Cui, M. Witbrock, M. A. Hasegawa-Johnson, and T. S. Huang · 2017
Cited alongside, same era.
Language modeling with gated convolutional networks
Y. Dauphin, A. Fan, M. Auli, and D. Grangier · 2017
Cited alongside, same era.
Fastgrnn: A fast, accurate, stable and tiny kilobyte sized gated recurrent neural network
A. Kusupati, M. Singh, K. Bhatia, A. Kumar, P. Jain, and M. Varma · 2018
Later among the works it cites.
Edge machine learning
Microsoft · 2018
Later among the works it cites.
When recurrent models don’t need to be recurrent
J. Miller and M. Hardt · 2018
Later among the works it cites.
Can recurrent neural networks warp time?
C. Tallec and Y. Ollivier · 2018
Later among the works it cites.
Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition
P. Warden · 2018
Later among the works it cites.
Stabilizing gradients for deep neural networks via efficient svd parameterization
J. Zhang, Q. Lei, and I. S. Dhillon · 2018
Later among the works it cites.
Stabilizing gradients for deep neural networks via efficient SVD parameterization
J. Zhang, Q. Lei, and I. S. Dhillon · 2018
Later among the works it cites.
AntisymmetricRNN: A dynamical system view on recurrent neural networks
B. Chang, M. Chen, E. Haber, and E. H. Chi · 2019
Closest in time.
Recurrent neural networks in the eye of differential equations
M. Y. Niu, L. Horesh, and I. Chuang · 2019
Closest in time.