Fetching the paper…
Reading the bibliography…
Truncated Backpropagation Through Time (truncated BPTT) is a widespread method for learning recurrent computational graphs.
Backpropagation through time: what does it do and how to do it
P. Werbos · 1990
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz · 1993
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Learning to forget: Continual prediction with LSTM
Felix A. Gers, Jürgen A. Schmidhuber, and Fred A. Cummins · 2000
Earlier work this paper cites.
A tutorial on training recurrent neural networks, covering BPPT, RTRL, EKF and the "echo state network" approach
Herbert Jaeger · 2005
Earlier work this paper cites.
Subword language modeling with neural networks
Tomás̆ Mikolov, Ilya Sutskever, Anoop Deoras, Le Hai-Son, Stefan Kombrink, and Jan C̆ernocký · 2012
Cited alongside, same era.
Generating sequences with recurrent neural networks
Alex Graves · 2013
Cited alongside, same era.
Training Recurrent Neural Networks
Ilya Sutskever · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Training recurrent networks online without backtracking
Yann Ollivier, Corentin Tallec, and Guillaume Charpiat · 2015
Later among the works it cites.
Tim Cooijmans, Nicolas Ballas, César Laurent, and Aaron C. Courville · 2016
Later among the works it cites.
Memory-efficient backpropagation through time
Audrunas Gruslys, Rémi Munos, Ivo Danihelka, Marc Lanctot, and Alex Graves · 2016
Later among the works it cites.
Unbiased online recurrent optimization
Corentin Tallec and Yann Ollivier · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…