Fetching the paper…
Reading the bibliography…
Stability is a fundamental property of dynamical systems, yet to this date it has had little bearing on the practice of recurrent neural networks.
Evaluation of spoken language systems: The atis domain
P. J. Price · 1990
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
M. P. Marcus, M. A. Marcinkiewicz, and B. Santorini · 1993
Earlier work this paper cites.
Absolute stability conditions for discrete-time recurrent neural networks
L. Jin, P. N. Nikiforuk, and M. M. Gupta · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Nonlinear Programming
D. P. Bertsekas · 1999
Earlier work this paper cites.
Harmonising chorales by probabilistic inference
M. Allan and C. Williams · 2005
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
R. Pascanu, T. Mikolov, and Y. Bengio · 2013
Earlier work this paper cites.
Using recurrent neural networks for slot filling in spoken language understanding
G. Mesnil, Y. Dauphin, K. Yao, Y. Bengio, L. Deng, D. Hakkani-Tur, X. He, L. Heck, G. Tur, D. Yu, et al · 2015
Earlier work this paper cites.
Unitary evolution recurrent neural networks
M. Arjovsky, A. Shah, and Y. Bengio · 2016
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
M. Hardt, B. Recht, and Y. Singer · 2016
Earlier work this paper cites.
Training input-output recurrent neural networks through spectral methods
H. Sedghi and A. Anandkumar · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
A. Van Den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Full-capacity unitary recurrent neural networks
S. Wisdom, T. Powers, J. Hershey, J. Le Roux, and L. Atlas · 2016
Cited alongside, same era.
Language modeling with gated convolutional networks
Y. N. Dauphin, A. Fan, M. Auli, and D. Grangier · 2017
Cited alongside, same era.
Convolutional sequence to sequence learning
J. Gehring, M. Auli, D. Grangier, D. Yarats, and Y. N. Dauphin · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Later among the works it cites.
On orthogonality and learning recurrent networks with long term dependencies
E. Vorontsov, C. Trabelsi, S. Kadoury, and C. Pal · 2017
Later among the works it cites.
An empirical evaluation of generic convolutional and recurrent networks for sequence modeling
S. Bai, J. Z. Kolter, and V. Koltun · 2018
Closest in time.
Gradient descent learns linear dynamical systems
M. Hardt, T. Ma, and B. Recht · 2018
Closest in time.
Kronecker recurrent units
C. Jose, M. Cisse, and F. Fleuret · 2018
Closest in time.
Fastgrnn: A fast, accurate, stable and tiny kilobyte sized gated recurrent neural network
A. Kusupati, M. Singh, K. Bhatia, A. Kumar, P. Jain, and M. Varma · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tunable efficient unitary neural networks (EUNN) and their application to rnns
L. Jing, Y. Shen, T. Dubcek, J. Peurifoy, S. Skirlo, Y. LeCun, M. Tegmark, and M. Soljačić · 2017
Cited alongside, same era.
A recurrent neural network without chaos
T. Laurent and J. von Brecht · 2017
Cited alongside, same era.
Pointer sentinel mixture models
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2017
Cited alongside, same era.
Efficient orthogonal parametrisation of recurrent neural networks using householder reflections
Z. Mhammedi, A. Hellicar, A. Rahman, and J. Bailey · 2017
Cited alongside, same era.
Non-asymptotic analysis of robust control from coarse-grained identification
S. Tu, R. Boczar, A. Packard, and B. Recht · 2017
Cited alongside, same era.
Closest in time.
Regularizing and Optimizing LSTM Language Models
S. Merity, N. S. Keskar, and R. Socher · 2018
Closest in time.
Stochastic gradient descent learns state equations with nonlinear activations
S. Oymak · 2018
Closest in time.
Stabilizing gradients for deep neural networks via efficient SVD parameterization
J. Zhang, Q. Lei, and I. Dhillon · 2018
Closest in time.