Fetching the paper…
Reading the bibliography…
Neural language models (NLM) have been shown to outperform conventional n-gram language models by a substantial margin in Automatic Speech Recognition (ASR) and other tasks.
W. Chen, D. Grangier, and M. Auli, “Strategies for training large vocabulary neural language models,” in
1985
Earlier work this paper cites.
M. McCloskey and N. J. Cohen, “Catastrophic interference in connectionist networks: The sequential learning problem,”
1989
Earlier work this paper cites.
R. Ratcliff, “Connectionist models of recognition memory: constraints imposed by learning and forgetting functions,”
1990
Earlier work this paper cites.
P. F. Brown, P. V. Desouza, R. L. Mercer, V. J. D. Pietra, and J. C. Lai, “Class-based n-gram models of natural language,”
1992
Earlier work this paper cites.
R. Kneser and H. Ney, “Improved backing-off for m-gram language modeling,” in
1995
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,”
1997
Earlier work this paper cites.
Y. Bengio, R. Ducharme, P. Vincent, and C. Janvin, “A neural probabilistic language model,”
2003
Earlier work this paper cites.
J. R. Bellegarda, “Statistical language model adaptation: review and perspectives,”
2004
Earlier work this paper cites.
F. Morin and Y. Bengio, “Hierarchical probabilistic neural network language model.” in
2005
Earlier work this paper cites.
H. Schwenk, “Continuous space language models,”
2007
Earlier work this paper cites.
T. Mikolov, M. Karafiát, L. Burget, J. â. Äernocký, and S. Khudanpur, “Recurrent neural network based language model,” in
2010
Earlier work this paper cites.
T. Mikolov, S. Kombrink, L. Burget, J. Cernocky, and S. Khudanpur, “Extensions of recurrent neural network language model,” in
2011
Cited alongside, same era.
A. Deoras, T. Mikolov, S. Kombrink, M. Karafiát, and S. Khudanpur, “Variational approximation of long-span language models for lvcsr,” in
2011
Cited alongside, same era.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in
2012
Cited alongside, same era.
A. Mnih and Y. W. Teh, “A fast and simple algorithm for training neural probabilistic language models,”
2012
Cited alongside, same era.
A. Vaswani, Y. Zhao, V. Fossum, and D. Chiang, “Decoding with large-scale neural language models improves translation,” in
2013
Cited alongside, same era.
X. Chen, X. Liu, M. J. F. Gales, and P. C. Woodland, “Recurrent neural network language model training with noise contrastive estimation for speech recognition,” in
2015
Later among the works it cites.
P. Aleksic, C. Allauzen, D. Elson, A. Kracun, D. M. Casado, and P. J. Moreno, “Improved recognition of contact names in voice commands,” in
2015
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
X. Chen, X. Liu, Y. Wang, M. J. F. Gales, and P. C. Woodland, “Efficient training and evaluation of recurrent neural network language models for automatic speech recognition,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Hori, Y. Kubo, and A. Nakamura, “Real-time one-pass decoding with recurrent neural network language model for speech recognition,” in
2014
Cited alongside, same era.
M. D. Zeiler and R. Fergus, “Visualizing and understanding convolutional networks,” in
2014
Cited alongside, same era.
Y. Shi, W.-Q. Zhang, M. Cai, and J. Liu, “Variance regularization of rnnlm for speech recognition,” in
2014
Cited alongside, same era.
J. Devlin, R. Zbib, Z. Huang, T. Lamar, R. M. Schwartz, and J. Makhoul, “Fast and robust neural network joint models for statistical machine translation,” in
2014
Cited alongside, same era.
2014
Cited alongside, same era.
J. Andreas and D. Klein, “When and why are log-linear models self-normalizing?” in
2015
Cited alongside, same era.
2016
Later among the works it cites.
B. Zoph, A. Vaswani, J. May, and K. Knight, “Simple, fast noise-contrastive estimation for large rnn vocabularies.” in
2016
Later among the works it cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Later among the works it cites.
A. Gandhe, A. Rastrow, and B. Hoffmeister, “Scalable language model adaptation for spoken dialogue systems,” in
2018
Later among the works it cites.
A. Raju, B. Hedayatnia, L. Liu, A. Gandhe, C. Khatri, A. Metallinou, A. Venkatesh, and A. Rastrow, “Contextual language model adaptation for conversational agents.” in
2018
Later among the works it cites.
J. Goldberger and O. Melamud, “Self-normalization properties of language modeling,”
2018
Later among the works it cites.
R. Krishnamoorthi, “Quantizing deep convolutional networks for efficient inference: A whitepaper.”
2018
Later among the works it cites.