Fetching the paper…
Reading the bibliography…
As the core component of Natural Language Processing (NLP) system, Language Model (LM) can provide word representation and probability indication of word sequences.
Perplexity—a measure of the difficulty of speech recognition tasks
F. Jelinek, R. L. Mercer, L. R. Bahl, and J. K. Baker · 1977
Earlier work this paper cites.
Can artificial neural networks learn language models?
W. Xu and A. Rudnicky · 2000
Earlier work this paper cites.
Quick training of probabilistic neural nets by importance sampling
Y. Bengio and J. Senecal · 2003
Earlier work this paper cites.
A neural probabilistic language model
Y. Bengio, R. Ducharme, P. Vincent, and C. Janvin · 2003
Earlier work this paper cites.
Hierarchical probabilistic neural network language model
F. Morin and Y. Bengio · 2005
Earlier work this paper cites.
Factored neural language models
A. Alexandrescu and K. Kirchhoff · 2006
Earlier work this paper cites.
Adaptive importance sampling to accelerate training of a neural probabilistic language model
Y. Bengio and J. Senecal · 2008
Earlier work this paper cites.
A scalable hierarchical distributed language model
A. Mnih and G. E. Hinton · 2008
Earlier work this paper cites.
Recurrent neural network based language model
T. Mikolov, M. Karafiát, L. Burget, J. Cernocký, and S. Khudanpur · 2010
Earlier work this paper cites.
Extensions of recurrent neural network language model
T. Mikolov, S. Kombrink, L. Burget, J. Cernocký, and S. Khudanpur · 2011
Earlier work this paper cites.
RNNLM - Recurrent Neural Network Language Modeling Toolkit
T. Mikolov, S. Kombrink, A. Deoras, and L. Burget · 2011
Earlier work this paper cites.
Context dependent recurrent neural network language model
T. Mikolov and G. Zweig · 2012
Earlier work this paper cites.
Subword Language Modeling With Neural Networks
T. Mikolov, I. Sutskever, A. Deoras, H. Le, and S. Kombrink · 2012
Earlier work this paper cites.
Impact of Word Classing on Recurrent Neural Network Language Model
Y. Si, Y. Guo, Y. Liu, J. Pan, and Y. Yan · 2012
Cited alongside, same era.
Neural network language model with cache
D. Soutner, Z. Loose, L. Müller, and A. Prazák · 2012
Cited alongside, same era.
LSTM neural networks for language modeling
M. Sundermeyer, R. Schlüter, and H. Ney · 2012
Cited alongside, same era.
Factored language model based on recurrent neural network
Y. Wu, X. Lu, H. Yamamoto, S. Matsuda, C. Hori, and H. Kashioka · 2012
Cited alongside, same era.
Hybrid speech recognition with deep bidirectional LSTM
A. Graves, N. Jaitly, and A. Mohamed · 2013
Cited alongside, same era.
Structured output layer neural network language models for speech recognition
H. S. Le, I. Oparin, A. Allauzen, J. Gauvain, and F. Yvon · 2013
Cited alongside, same era.
Larger-context language modelling
T. Wang and K. Cho · 2015
Later among the works it cites.
A neural knowledge language model
S. Ahn, H. Choi, T. Pärnamaa, and Y. Bengio · 2016
Later among the works it cites.
Theanolm - an extensible toolkit for neural network language modeling
S. Enarvi and M. Kurimo · 2016
Later among the works it cites.
Improving neural language models with a continuous cache
E. Grave, A. Joulin, and N. Usunier · 2016
Later among the works it cites.
Character-level language modeling with hierarchical recurrent neural networks
K. Hwang and W. Sung · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Cited alongside, same era.
CSLM - a modular open-source continuous space language modeling toolkit
H. Schwenk · 2013
Cited alongside, same era.
Speed regularization and optimality in word classing
G. Zweig and K. Makarychev · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
Cache based recurrent neural network language model inference for first pass speech recognition
Z. Huang, G. Zweig, and B. Dumoulin · 2014
Cited alongside, same era.
Syntactic and semantic features for code-switching factored language models
H. Adel, N. T. Vu, K. Kirchhoff, D. Telaar, and T. Schultz · 2015
Cited alongside, same era.
Coherent dialogue with attention-based language models
H. Mei, M. Bansal, and M. R. Walter · 2016
Later among the works it cites.
Gated word-character recurrent language model
Y. Miyamoto and Ky. Cho · 2016
Later among the works it cites.
Recurrent memory networks for language modeling
K. M. Tran, A. Bisazza, and C. Monz · 2016
Later among the works it cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Later among the works it cites.
Character-word LSTM language models
L. Verwimp, J. Pelemans, H. Van hamme, and P. Wambacq · 2017
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova · 2018
Later among the works it cites.
Deep contextualized word representations
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer · 2018
Later among the works it cites.
Improving Language Understanding by Generative Pre-Training
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever · 2018
Later among the works it cites.