Fetching the paper…
Reading the bibliography…
Transformers have a potential of learning longer-term dependency, but are limited by a fixed-length context in the setting of language modeling.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies
Sepp Hochreiter, Yoshua Bengio, Paolo Frasconi, Jürgen Schmidhuber, et al. 2001 · 2001
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin. 2003 · 2003
Earlier work this paper cites.
Hierarchical probabilistic neural network language model
Frederic Morin and Yoshua Bengio. 2005 · 2005
Earlier work this paper cites.
Large text compression benchmark
MultiMedia LLC. 2009 · 2009
Earlier work this paper cites.
Recurrent neural network based language model
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur. 2010 · 2010
Earlier work this paper cites.
Context dependent recurrent neural network language model
Tomas Mikolov and Geoffrey Zweig. 2012 · 2012
Earlier work this paper cites.
Understanding the exploding gradient problem
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. 2012 · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson. 2013 · 2013
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Alex Graves. 2013 · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka. 2014 · 2014
Earlier work this paper cites.
Jan Koutnik, Klaus Greff, Faustino Gomez, and Juergen Schmidhuber. 2014 · 2014
Earlier work this paper cites.
Learning longer memory in recurrent neural networks
Tomas Mikolov, Armand Joulin, Sumit Chopra, Michael Mathieu, and Marc’Aurelio Ranzato. 2014 · 2014
Earlier work this paper cites.
Skip-gram language modeling using sparse non-negative matrix probability estimation
Noam Shazeer, Joris Pelemans, and Ciprian Chelba. 2014 · 2014
Earlier work this paper cites.
Jason Weston, Sumit Chopra, and Antoine Bordes. 2014 · 2014
Earlier work this paper cites.
Recurrent neural network regularization
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals. 2014 · 2014
Earlier work this paper cites.
Semi-supervised sequence learning
Andrew M Dai and Quoc V Le. 2015 · 2015
Earlier work this paper cites.
Document context language models
Yangfeng Ji, Trevor Cohn, Lingpeng Kong, Chris Dyer, and Jacob Eisenstein. 2015 · 2015
Earlier work this paper cites.
A simple way to initialize recurrent networks of rectified linear units
Quoc V Le, Navdeep Jaitly, and Geoffrey E Hinton. 2015 · 2015
Earlier work this paper cites.
Larger-context language modelling
Tian Wang and Kyunghyun Cho. 2015 · 2015
Earlier work this paper cites.
Hierarchical multiscale recurrent neural networks
Junyoung Chung, Sungjin Ahn, and Yoshua Bengio. 2016 · 2016
Cited alongside, same era.
Tim Cooijmans, Nicolas Ballas, César Laurent, Çağlar Gülçehre, and Aaron Courville. 2016 · 2016
Cited alongside, same era.
Language modeling with gated convolutional networks
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier. 2016 · 2016
Cited alongside, same era.
Topicrnn: A recurrent neural network with long-range semantic dependency
Adji B Dieng, Chong Wang, Jianfeng Gao, and John Paisley. 2016 · 2016
Cited alongside, same era.
A theoretically grounded application of dropout in recurrent neural networks
Topic compositional neural language model
Wenlin Wang, Zhe Gan, Wenqi Wang, Dinghan Shen, Jiaji Huang, Wei Ping, Sanjeev Satheesh, and Lawrence Carin. 2017 · 2017
Later among the works it cites.
Breaking the softmax bottleneck: A high-rank rnn language model
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W Cohen. 2017 · 2017
Later among the works it cites.
Character-level language modeling with deeper self-attention
Rami Al-Rfou, Dokook Choe, Noah Constant, Mandy Guo, and Llion Jones. 2018 · 2018
Later among the works it cites.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli. 2018 · 2018
Later among the works it cites.
An empirical evaluation of generic convolutional and recurrent networks for sequence modeling
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yarin Gal and Zoubin Ghahramani. 2016 · 2016
Cited alongside, same era.
David Ha, Andrew Dai, and Quoc V Le. 2016 · 2016
Cited alongside, same era.
Tying word vectors and word classifiers: A loss framework for language modeling
Hakan Inan, Khashayar Khosravi, and Richard Socher. 2016 · 2016
Cited alongside, same era.
Exploring the limits of language modeling
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. 2016 · 2016
Cited alongside, same era.
Neural machine translation in linear time
Nal Kalchbrenner, Lasse Espeholt, Karen Simonyan, Aaron van den Oord, Alex Graves, and Koray Kavukcuoglu. 2016 · 2016
Cited alongside, same era.
Multiplicative lstm for sequence modelling
Ben Krause, Liang Lu, Iain Murray, and Steve Renals. 2016 · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016 · 2016
Cited alongside, same era.
Using the output embedding to improve language models
Ofir Press and Lior Wolf. 2016 · 2016
Cited alongside, same era.
Shaojie Bai, J Zico Kolter, and Vladlen Koltun. 2018 · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
An improved relative self-attention mechanism for transformer with application to music generation
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Noam Shazeer, Curtis Hawthorne, Andrew M Dai, Matthew D Hoffman, and Douglas Eck. 2018 · 2018
Later among the works it cites.
Sigsoftmax: Reanalysis of the softmax bottleneck
Sekitoshi Kanai, Yasuhiro Fujiwara, Yuki Yamanaka, and Shuichi Adachi. 2018 · 2018
Later among the works it cites.
Sparse attentive backtracking: Temporal credit assignment through reminding
Nan Rosemary Ke, Anirudh Goyal ALIAS PARTH GOYAL, Olexa Bilaniuk, Jonathan Binas, Michael C Mozer, Chris Pal, and Yoshua Bengio. 2018 · 2018
Later among the works it cites.
Sharp nearby, fuzzy far away: How neural language models use context
Urvashi Khandelwal, He He, Peng Qi, and Dan Jurafsky. 2018 · 2018
Later among the works it cites.
Independently recurrent neural network (indrnn): Building a longer and deeper rnn
Shuai Li, Wanqing Li, Chris Cook, Ce Zhu, and Yanbo Gao. 2018 · 2018
Later among the works it cites.
Darts: Differentiable architecture search
Hanxiao Liu, Karen Simonyan, and Yiming Yang. 2018 · 2018
Later among the works it cites.
Gábor Melis, Charles Blundell, Tomáš Kočiskỳ, Karl Moritz Hermann, Chris Dyer, and Phil Blunsom. 2018 · 2018
Later among the works it cites.
An analysis of neural language modeling at multiple scales
Stephen Merity, Nitish Shirish Keskar, and Richard Socher. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Efficient neural architecture search via parameter sharing
Hieu Pham, Melody Y Guan, Barret Zoph, Quoc V Le, and Jeff Dean. 2018 · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
Fast parametric learning with activation memorization
Jack W Rae, Chris Dyer, Peter Dayan, and Timothy P Lillicrap. 2018 · 2018
Later among the works it cites.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2018
Later among the works it cites.
Mesh-tensorflow: Deep learning for supercomputers
Noam Shazeer, Youlong Cheng, Niki Parmar, Dustin Tran, Ashish Vaswani, Penporn Koanantakool, Peter Hawkins, HyoukJoong Lee, Mingsheng Hong, Cliff Young, et al. 2018 · 2018
Later among the works it cites.
Learning longer-term dependencies in rnns with auxiliary losses
Trieu H Trinh, Andrew M Dai, Thang Luong, and Quoc V Le. 2018 · 2018
Later among the works it cites.