Fetching the paper…
Reading the bibliography…
It is often the case that the best performing language model is an ensemble of a neural language model with n-grams.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
A maximum likelihood approach to continuous speech recognition
Lalit R Bahl, Frederick Jelinek, and Robert L Mercer. 1990 · 1990
Earlier work this paper cites.
Adaptive mixtures of local experts
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton. 1991 · 1991
Earlier work this paper cites.
The mathematics of statistical machine translation: Parameter estimation
Peter F Brown, Vincent J Della Pietra, Stephen A Della Pietra, and Robert L Mercer. 1993 · 1993
Earlier work this paper cites.
On the dynamic adaptation of stochastic language models
Reinhard Kneser and Volker Steinbiss. 1993 · 1993
Earlier work this paper cites.
Improved backing-off for m-gram language modeling
Reinhard Kneser and Hermann Ney. 1995 · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jurgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
An empirical study of smoothing techniques for language modeling
Stanley F Chen and Joshua Goodman. 1999 · 1999
Cited alongside, same era.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin. 2003 · 2003
Cited alongside, same era.
On using very large target vocabulary for neural machine translation
Sébastien Jean, Kyunghyun Cho, Roland Memisevic, and Yoshua Bengio. 2014 · 2007
Cited alongside, same era.
Recurrent neural network based language model
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur. 2010 · 2010
Cited alongside, same era.
Kenlm: Faster and smaller language model queries
Kenneth Heafield. 2011 · 2011
Cited alongside, same era.
Strategies for training large scale neural network language models
Tomáš Mikolov, Anoop Deoras, Daniel Povey, Lukáš Burget, and Jan Černockỳ. 2011 · 2011
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Later among the works it cites.
Language modeling with gated convolutional networks
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier. 2016 · 2016
Later among the works it cites.
Efficient softmax approximation for GPUs
Edouard Grave, Armand Joulin, Moustapha Cissé, David Grangier, and Hervé Jégou. 2016 · 2016
Later among the works it cites.
Exploring the limits of language modeling
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. 2016 · 2016
Later among the works it cites.
Generalizing and hybridizing count-based and neural language models
Graham Neubig and Chris Dyer. 2016 · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson. 2013 · 2013
Cited alongside, same era.
Later among the works it cites.
A theoretical model for n-gram distribution in big data corpora
Joaquim F Silva, Carlos Goncalves, and Jose C Cunha. 2016 · 2016
Later among the works it cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017 · 2017
Later among the works it cites.