Fetching the paper…
Reading the bibliography…
We reformulate the problem of encoding a multi-scale representation of a sequence in a language model by casting it in a continuous learning framework.
Learning Complex, Extended Sequences Using the Principle of History Compression
J. Schmidhuber. 1992 · 1992
Earlier work this paper cites.
Uniqueness of the weights for minimal feedforward nets with a given input-output map
Héctor J. Sussmann. 1992 · 1992
Earlier work this paper cites.
Hierarchical Recurrent Neural Networks for Long-term Dependencies
Salah El Hihi and Yoshua Bengio. 1995 · 1995
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Neural Networks with Dynamic Synapses
Misha Tsodyks, Klaus Pawelzik, and Henry Markram. 1998 · 1998
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
Robert M. French. 1999 · 1999
Earlier work this paper cites.
Synaptic computation
L. F. Abbott and Wade G. Regehr. 2004 · 2004
Earlier work this paper cites.
Persistent activity in neural networks with dynamic synapses
Omri Barak and Misha Tsodyks. 2007 · 2007
Earlier work this paper cites.
Recurrent neural network based language model , volume 2
Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernocký, and Sanjeev Khudanpur. 2010 · 2010
Earlier work this paper cites.
Deep neural network language models
Ebru Arisoy, Tara N. Sainath, Brian Kingsbury, and Bhuvana Ramabhadran. 2012 · 2012
Earlier work this paper cites.
Continuous space translation models for phrase-based statistical machine translation
Holger Schwenk. 2012 · 2012
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Cited alongside, same era.
Alex Graves, Greg Wayne, and Ivo Danihelka. 2014 · 2014
Cited alongside, same era.
A Clockwork RNN
Jan Koutník, Klaus Greff, Faustino Gomez, and Jürgen Schmidhuber. 2014 · 2014
Cited alongside, same era.
A neural attention model for abstractive sentence summarization
Alexander M. Rush, Sumit Chopra, and Jason Weston. 2015 · 2015
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W. Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando de Freitas. 2016 · 2016
Pointer Sentinel Mixture Models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016 · 2016
Later among the works it cites.
Abstractive text summarization using sequence-to-sequence rnns and beyond
Ramesh Nallapati, Bowen Zhou, Caglar Gulcehre, and Bing Xiang. 2016 · 2016
Later among the works it cites.
Optimization as a Model for Few-Shot Learning
Sachin Ravi and Hugo Larochelle. 2016 · 2016
Later among the works it cites.
Recurrent Memory Networks for Language Modeling
Ke Tran, Arianna Bisazza, and Christof Monz. 2016 · 2016
Later among the works it cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell. 2017 · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Using Fast Weights to Attend to the Recent Past
Jimmy Ba, Geoffrey Hinton, Volodymyr Mnih, Joel Z. Leibo, and Catalin Ionescu. 2016 · 2016
Cited alongside, same era.
Hierarchical Multiscale Recurrent Neural Networks
Junyoung Chung, Sungjin Ahn, and Yoshua Bengio. 2016 · 2016
Cited alongside, same era.
Frustratingly Short Attention Spans in Neural Language Modeling
Michał Daniluk, Tim Rocktäschel, Johannes Welbl, and Sebastian Riedel. 2016 · 2016
Cited alongside, same era.
Memory-Efficient Backpropagation Through Time
Audrūnas Gruslys, Remi Munos, Ivo Danihelka, Marc Lanctot, and Alex Graves. 2016 · 2016
Cited alongside, same era.
David Ha, Andrew Dai, and Quoc V. Le. 2016 · 2016
Cited alongside, same era.
Ke Li and Jitendra Malik. 2016 · 2016
Cited alongside, same era.
Later among the works it cites.
Dynamic Evaluation of Neural Sequence Models
Ben Krause, Emmanuel Kahembwe, Iain Murray, and Steve Renals. 2017 · 2017
Later among the works it cites.
Regularizing and Optimizing LSTM Language Models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher. 2017 · 2017
Later among the works it cites.
Learning to Generate Reviews and Discovering Sentiment
Alec Radford, Rafal Jozefowicz, and Ilya Sutskever. 2017 · 2017
Later among the works it cites.
Fine-tuned Language Models for Text Classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Closest in time.
Joel Ruben Antony Moniz and David Krueger. 2018 · 2018
Closest in time.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Closest in time.