Fetching the paper…
Reading the bibliography…
Current methods for training recurrent neural networks are based on backpropagation through time, which requires storing a complete history of network states, and prohibits updating the weights `online' (after every timestep).
On the variance of unbiased online recurrent optimization
Tim Cooijmans and James Martens · 1902
Earlier work this paper cites.
Erich Elsen, Marat Dukhan, Trevor Gale, and Karen Simonyan · 1911
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
R. J. Williams and D. Zipser · 1989
Earlier work this paper cites.
An efficient gradient-based algorithm for on-line training of recurrent network trajectories
Ronald J. Williams and Jing Peng · 1990
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Sparse connection and pruning in large dynamic artificial neural networks
Nikko Ström · 1997
Earlier work this paper cites.
Algorithm 799: Revolve: An implementation of checkpointing for the reverse or adjoint mode of computational differentiation
Andreas Griewank and Andrea Walther · 2000
Earlier work this paper cites.
Learning phrase representations using RNN encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2014
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Deep speech 2: End-to-end speech recognition in english and mandarin
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, and et al · 2016
Earlier work this paper cites.
Hybrid computing using a neural network with dynamic external memory
Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka Grabska-Barwińska, Sergio Gómez Colmenarejo, Edward Grefenstette, Tiago Ramalho, John Agapiou, AdriàPuigdomènech Badia, Karl Moritz Hermann, Yori Zwols, Georg Ostrovski, Adam Cain, Helen King, Christopher Summerfield, Phil Blunsom, Koray Kavukcuoglu, and Demis Hassabis · 2016
Earlier work this paper cites.
Memory-efficient backpropagation through time
Audrunas Gruslys, Remi Munos, Ivo Danihelka, Marc Lanctot, and Alex Graves · 2016
Cited alongside, same era.
Scaling memory-augmented neural networks with sparse reads and writes
Jack W Rae, Jonathan J Hunt, Tim Harley, Ivo Danihelka, Andrew Senior, Greg Wayne, Alex Graves, and Timothy P Lillicrap · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Cited alongside, same era.
Exploring sparsity in recurrent neural networks
Sharan Narang, Gregory F. Diamos, Shubho Sengupta, and Erich Elsen · 2017
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, and Skye Wanderman-Milne · 2018
Cited alongside, same era.
The best of both worlds: Combining recent advances in neural machine translation
Unbiased online recurrent optimization
Corentin Tallec and Yann Ollivier · 2018
Later among the works it cites.
To Prune, or Not to Prune: Exploring the Efficacy of Pruning for Model Compression
Michael Zhu and Suyog Gupta · 2018
Later among the works it cites.
Optimal kronecker-sum approximation of real time recurrent learning
Frederik Benzing, Marcelo Matheus Gauy, Asier Mujika, Anders Martinsson, and Angelika Steger · 2019
Later among the works it cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov · 2019
Later among the works it cites.
Rigging the Lottery: Making All Tickets Winners, 2019
Utku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro, and Erich Elsen · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George Foster, Llion Jones, Mike Schuster, Noam Shazeer, Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Zhifeng Chen, Yonghui Wu, and Macduff Hughes · 2018
Cited alongside, same era.
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser · 2018
Cited alongside, same era.
IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymir Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Cited alongside, same era.
Efficient Neural Audio Synthesis
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aäron van den Oord, Sander Dieleman, and Koray Kavukcuoglu · 2018
Cited alongside, same era.
Approximating real-time recurrent learning with random kronecker factors
Asier Mujika, Florian Meier, and Angelika Steger · 2018
Cited alongside, same era.
Jonathan Frankle and Michael Carbin · 2019
Later among the works it cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, Radu Soricut, Google Research, and Mailton de Carvalho · 2019
Later among the works it cites.
Optimizing rnns with differentiable graphs
Jesse Engel · 2020
Closest in time.
Haiku: Sonnet for JAX, 2020
Tom Hennigan, Trevor Cai, Tamara Norman, and Igor Babuschkin · 2020
Closest in time.
Local online learning in recurrent networks with random feedback
James M Murray · 2050
Closest in time.