Fetching the paper…
Reading the bibliography…
Recurrent neural networks (RNNs) are widely used to model sequential data but their non-linear dependencies between sequence elements prevent parallelizing training over sequence length.
Parallel prefix computation
R. E. Ladner and M. J. Fischer · 1980
Earlier work this paper cites.
Prefix sums and their applications
G. E. Blelloch · 1990
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
The mnist database of handwritten digits, 1998
Y. LeCun, C. Cortes, and C. J. Burges · 1998
Earlier work this paper cites.
PhysioBank, PhysioToolkit, and PhysioNet: Components of a New Research Resource for Complex Physiologic Signals
Goldberger et al · 2000
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Cited alongside, same era.
Deep speech 2: End-to-end speech recognition in english and mandarin
D. Amodei, R. Anubhai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, J. Chen, M. Chrzanowski, A. Coates, G. Diamos, et al · 2015
Cited alongside, same era.
Deep recurrent q-learning for partially observable mdps
M. Hausknecht and P. Stone · 2015
Cited alongside, same era.
Converting static image datasets to spiking neuromorphic datasets using saccades
G. Orchard, A. Jayawant, G. Cohen, and N. Thakor · 2015
Cited alongside, same era.
Persistent rnns: Stashing recurrent weights on-chip
G. Diamos, S. Sengupta, B. Catanzaro, M. Chrzanowski, A. Coates, E. Elsen, J. Engel, A. Hannun, and S. Satheesh · 2016
Later among the works it cites.
Neural machine translation in linear time
N. Kalchbrenner, L. Espeholt, K. Simonyan, A. v. d. Oord, A. Graves, and K. Kavukcuoglu · 2016
Later among the works it cites.
Wavenet: A generative model for raw audio
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu · 2016
Later among the works it cites.
Quasi-recurrent neural networks
J. Bradbury, S. Merity, C. Xiong, and R. Socher · 2017
Closest in time.
Convolutional sequence to sequence learning
J. Gehring, M. Auli, D. Grangier, D. Yarats, and Y. N. Dauphin · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, et al · 2016
Cited alongside, same era.
Strongly-typed recurrent neural networks
D. Balduzzi and M. Ghifary · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2017
Closest in time.
T. Lei, Y. Zhang, · 2017
Closest in time.