Fetching the paper…
Reading the bibliography…
The pre-dominant approach to language modeling to date is based on recurrent neural networks.
Improved backing-off for m-gram language modeling
Kneser, Reinhard and Ney, Hermann · 1995
Earlier work this paper cites.
Convolutional networks for images, speech, and time series
LeCun, Yann and Bengio, Yoshua · 1995
Earlier work this paper cites.
An empirical study of smoothing techniques for language modeling
Chen, Stanley F and Goodman, Joshua · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Foundations of statistical natural language processing, 1999
Manning, Christopher D and Schütze, Hinrich · 1999
Earlier work this paper cites.
The syntactic process
Steedman, Mark · 2002
Earlier work this paper cites.
A neural probabilistic language model
Bengio, Yoshua, Ducharme, Réjean, Vincent, Pascal, and Jauvin, Christian · 2003
Earlier work this paper cites.
Hierarchical probabilistic neural network language model
Morin, Frederic and Bengio, Yoshua · 2005
Earlier work this paper cites.
Three new graphical models for statistical language modelling
Mnih, Andriy and Hinton, Geoffrey · 2007
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, Xavier and Bengio, Yoshua · 2010
Earlier work this paper cites.
Statistical Machine Translation
Koehn, Philipp · 2010
Cited alongside, same era.
Recurrent Neural Network based Language Model
Mikolov, Tomáš, Martin, Karafiát, Burget, Lukáš, Cernocký, Jan, and Khudanpur, Sanjeev · 2010
Cited alongside, same era.
Torch7: A Matlab-like Environment for Machine Learning
Collobert, Ronan, Kavukcuoglu, Koray, and Farabet, Clement · 2011
Cited alongside, same era.
One billion word benchmark for measuring progress in statistical language modeling
Chelba, Ciprian, Mikolov, Tomas, Schuster, Mike, Ge, Qi, Brants, Thorsten, Koehn, Phillipp, and Robinson, Tony · 2013
Cited alongside, same era.
On the difficulty of training recurrent neural networks
Pascanu, Razvan, Mikolov, Tomas, and Bengio, Yoshua · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Blackout: Speeding up recurrent neural network language models with very large vocabularies
Ji, Shihao, Vishwanathan, SVN, Satish, Nadathur, Anderson, Michael J, and Dubey, Pradeep · 2015
Later among the works it cites.
gen cnn: A convolutional architecture for word sequence prediction
Wang, Mingxuan, Lu, Zhengdong, Li, Hang, Jiang, Wenbin, and Liu, Qun · 2015
Later among the works it cites.
Strategies for training large vocabulary neural language models
Chen, Wenlin, Grangier, David, and Auli, Michael · 2016
Closest in time.
Exploring the limits of language modeling
Jozefowicz, Rafal, Vinyals, Oriol, Schuster, Mike, Shazeer, Noam, and Wu, Yonghui · 2016
Closest in time.
Neural Machine Translation in Linear Time
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sutskever, Ilya, Martens, James, Dahl, George E, and Hinton, Geoffrey E · 2013
Cited alongside, same era.
Skip-gram language modeling using sparse non-negative matrix probability estimation
Shazeer, Noam, Pelemans, Joris, and Chelba, Ciprian · 2014
Cited alongside, same era.
Automatic Speech Recognition: A Deep Learning Approach
Yu, Dong and Deng, Li · 2014
Cited alongside, same era.
Predicting distributions with linearizing belief networks
Dauphin, Yann N and Grangier, David · 2015
Cited alongside, same era.
Efficient softmax approximation for GPUs
Grave, E., Joulin, A., Cissé, M., Grangier, D., and Jégou, H
Cited in the paper.
Improving Neural Language Models with a Continuous Cache
Grave, E., Joulin, A., and Usunier, N
Cited in the paper.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, Michael and Hyvärinen, Aapo
Cited in the paper.
Kalchbrenner, Nal, Espeholt, Lasse, Simonyan, Karen, van den Oord, Aaron, Graves, Alex, and Kavukcuoglu, Koray · 2016
Closest in time.
Pointer Sentinel Mixture Models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Closest in time.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, Tim and Kingma, Diederik P · 2016
Closest in time.
Factorization tricks for LSTM networks
Kuchaiev, Oleksii and Ginsburg, Boris · 2017
Closest in time.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, Noam, Mirhoseini, Azalia, Maziarz, Krzysztof, Davis, Andy, Le, Quoc V., Hinton, Geoffrey E., and Dean, Jeff · 2017
Closest in time.