Fetching the paper…
Reading the bibliography…
Truncated backpropagation through time (TBPTT) is a popular method for learning in recurrent neural networks (RNNs) that saves computation and memory at the cost of bias by truncating backpropagation after a fixed number of lags.
Pseudogradient adaptation and training algorithms
BT Poljak and Ya Z Tsypkin · 1973
Earlier work this paper cites.
Parallel and Distributed Computation: Numerical Methods , volume 23
Dimitri P Bertsekas and John N Tsitsiklis · 1989
Earlier work this paper cites.
Backpropagation through time: What it does and how to do it
Paul J Werbos et al · 1990
Earlier work this paper cites.
Building a large annotated corpus of english: The Penn treebank
Mitchell Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz · 1993
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, Paolo Frasconi, et al · 1994
Earlier work this paper cites.
Gradient-based learning algorithms for recurrent networks and their computational complexity
Ronald J Williams and David Zipser · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Recurrent neural network based language model
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur · 2010
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course , volume 87
Yurii Nesterov · 2013
Cited alongside, same era.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Cited alongside, same era.
Training Recurrent Neural Networks
Ilya Sutskever · 2013
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Cited alongside, same era.
Unitary evolution recurrent neural networks
Martin Arjovsky, Amar Shah, and Yoshua Bengio · 2016
Cited alongside, same era.
Recurrent marked temporal point processes: Embedding event history to vector
Nan Du, Hanjun Dai, Rakshit Trivedi, Utkarsh Upadhyay, Manuel Gomez-Rodriguez, and Le Song · 2016
Tunable efficient unitary neural networks (EUNN) and their application to RNNs
Li Jing, Yichen Shen, Tena Dubcek, John Peurifoy, Scott Skirlo, Yann LeCun, Max Tegmark, and Marin Soljačić · 2017
Later among the works it cites.
A recurrent neural network without chaos
Thomas Laurent and James von Brecht · 2017
Later among the works it cites.
Tao Lei, Yu Zhang, and Yoav Artzi · 2017
Later among the works it cites.
On orthogonality and learning recurrent networks with long term dependencies
Eugene Vorontsov, Chiheb Trabelsi, Samuel Kadoury, and Chris Pal · 2017
Later among the works it cites.
Factorized recurrent neural architectures for longer range dependence
Francois Belletti, Alex Beutel, Sagar Jain, and Ed Chi · 2018
Later among the works it cites.
Stochastic gradient descent with biased but consistent gradient estimators
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Cited alongside, same era.
Stochastic variance reduction for nonconvex optimization
Sashank J Reddi, Ahmed Hefny, Suvrit Sra, Barnabas Poczos, and Alex Smola · 2016
Cited alongside, same era.
Jie Chen and Ronny Luss · 2018
Later among the works it cites.
Approximating real-time recurrent learning with random Kronecker factors
Asier Mujika, Florian Meier, and Angelika Steger · 2018
Later among the works it cites.
Stable recurrent models
John Miller and Moritz Hardt · 2019
Closest in time.