Fetching the paper…
Reading the bibliography…
Language models are the foundation of current neural network-based models for natural language understanding and generation.
A Mathematical Theory of Communication
C. E. Shannon. 1948 · 1948
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
A new algorithm for data compression
Philip Gage. 1994 · 1994
Earlier work this paper cites.
Improved backing-off for M-gram language modeling
Reinhard Kneser and Hermann Ney. 1995 · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
An empirical study of smoothing techniques for language modeling
Stanley F Chen and Joshua Goodman. 1999 · 1999
Earlier work this paper cites.
Searching the web by voice
Alexander Franz and Brian Milch. 2002 · 2002
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin. 2003 · 2003
Earlier work this paper cites.
Exploiting cross-linguistic similarities in Zulu and Xhosa computational morphology
Laurette Pretorius and Sonja Bosch. 2009 · 2009
Earlier work this paper cites.
Census 2011, Census in Brief
Statistics South Africa. 2012 · 2011
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
Regularization of neural networks using dropconnect
Li Wan, Matthew Zeiler, Sixin Zhang, Yann Le Cun, and Rob Fergus. 2013 · 2013
Earlier work this paper cites.
Developing text resources for ten south African languages
Roald Eiselen and Martin Puttkammer. 2014 · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
Recurrent Neural Network Regularization
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals. 2015 · 2015
Cited alongside, same era.
Assessing the impact of vocabulary similarity on multilingual information retrieval for bantu languages
Catherine Chavula and Hussein Suleman. 2016 · 2016
Cited alongside, same era.
A theoretically grounded application of dropout in recurrent neural networks
Yarin Gal and Zoubin Ghahramani. 2016 · 2016
Cited alongside, same era.
Revisiting Activation Regularization for Language RNNs
Stephen Merity, Bryan Mccann, and Richard Socher. 2017 · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
Generating wikipedia by summarizing long sequences
Peter J. Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer. 2018 · 2018
Later among the works it cites.
Direct output connection for a high-rank language model
Sho Takase, Jun Suzuki, and Masaaki Nagata. 2018 · 2018
Later among the works it cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ankit Kumar, Ozan Irsoy, Peter Ondruska, Mohit Iyyer, James Bradbury, Ishaan Gulrajani, Victor Zhong, Romain Paulus, and Richard Socher. 2016 · 2016
Cited alongside, same era.
The effects of a corpus on isizulu spellcheckers based on n-grams
B. Ndaba, H. Suleman, C. M. Keet, and L. Khumalo. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016 · 2016
Cited alongside, same era.
Quasi-Recurrent Neural Networks
James Bradbury, Stephen Merity, Caiming Xiong, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Improving neural language models with a continuous cache
Edouard Grave, Armand Joulin, and Nicolas Usunier. 2017 · 2017
Cited alongside, same era.
An Analysis of Neural Language Modeling at Multiple Scales
Stephen Merity, Nitish Shirish Keskar, and Richard Socher. 2018a
Cited in the paper.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Spell once, summon anywhere: A two-level open-vocabulary language model
Sebastian J Mielke and Jason Eisner. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
Later among the works it cites.
Scaling hidden Markov language models
Justin Chiu and Alexander Rush. 2020 · 2020
Later among the works it cites.
N-gram Language Models
Daniel Jurafsky and James H Martin. 2020 · 2020
Later among the works it cites.