Fetching the paper…
Reading the bibliography…
We investigate the effective memory depth of RNN models by using them for $n$-gram language model (LM) smoothing.
Estimation of probabilities from sparse data for the language model component of a speech recognizer
S. Katz · 1987
Earlier work this paper cites.
Improved backing-off for m-gram language modeling
R. Kneser and H. Ney · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Information Extraction From Speech And Text , chapter 8, pp. 141–142
Frederick Jelinek · 1997
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, and Pascal Vincent · 2001
Earlier work this paper cites.
A bit of progress in language modeling, extended version
Joshua Goodman · 2001
Cited alongside, same era.
Large language models in machine translation
T. Brants, A. C. Popat, P. Xu, F. J. Och, and J. Dean · 2007
Cited alongside, same era.
Recurrent neural network based language model
Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernockỳ, and Sanjeev Khudanpur · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Cited alongside, same era.
Understanding the exploding gradient problem
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2012
Cited alongside, same era.
Recurrent neural networks tutorial (language modeling), 2015a
M. Abadi and et al
Cited in the paper.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015b
M. Abadi and et al
Cited in the paper.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, and Phillipp Koehn · 2013
Later among the works it cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Later among the works it cites.
Single machine implementation of LSTM language model on One Billion Words benchmark using synchronized gradient updates, 2016
Rafal Józefowicz · 2016
Later among the works it cites.
Exploring the limits of language modeling
Rafal Józefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu · 2016
Later among the works it cites.
Sparse non-negative matrix language modeling
Joris Pelemans, Noam Shazeer, and Ciprian Chelba · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…