Fetching the paper…
Reading the bibliography…
Despite all the impressive advances of recurrent neural networks, sequential data is still in need of better modelling.
Learning representations by back-propagating errors
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
R. J. Williams and D. Zipser · 1989
Earlier work this paper cites.
An efficient gradient-based algorithm for on-line training of recurrent network trajectories
R. J. Williams and J. Peng · 1990
Earlier work this paper cites.
A method for improving the real-time recurrent learning algorithm
T. Catfolis · 1993
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
M. P. Marcus, M. A. Marcinkiewicz, and B. Santorini · 1993
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P. Frasconi · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Cited alongside, same era.
The “echo state” approach to analysing and training recurrent neural networks-with an erratum note
H. Jaeger · 2001
Cited alongside, same era.
Real-time computing without stable states: A new framework for neural computation based on perturbations
W. Maass, T. Natschläger, and H. Markram · 2002
Cited alongside, same era.
Reservoir computing approaches to recurrent neural network training
M. Lukoševičius and H. Jaeger · 2009
Cited alongside, same era.
Subword language modeling with neural networks
T. Mikolov, I. Sutskever, A. Deoras, H.-S. Le, S. Kombrink, and J. Cernocky · 2012
Cited alongside, same era.
Training recurrent networks online without backtracking
Y. Ollivier, C. Tallec, and G. Charpiat · 2015
Later among the works it cites.
Tensorflow: A system for large-scale machine learning
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, et al · 2016
Later among the works it cites.
Decoupled neural interfaces using synthetic gradients
M. Jaderberg, W. M. Czarnecki, S. Osindero, O. Vinyals, A. Graves, D. Silver, and K. Kavukcuoglu · 2016
Later among the works it cites.
J. G. Zilly, R. K. Srivastava, J. Koutník, and J. Schmidhuber · 2016
Later among the works it cites.
On the state of the art of evaluation in neural language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
Unbiased online recurrent optimization
C. Tallec and Y. Ollivier
Cited in the paper.
Unbiasing truncated backpropagation through time
C. Tallec and Y. Ollivier
Cited in the paper.
G. Melis, C. Dyer, and P. Blunsom · 2017
Later among the works it cites.
An analysis of neural language modeling at multiple scales
S. Merity, N. S. Keskar, and R. Socher · 2018
Closest in time.