Fetching the paper…
Reading the bibliography…
Recurrent neural networks have proved to be an effective method for statistical language modeling.
Interpolated estimation of Markov source parameters from sparse data,
F. Jelinek, R. L. Mercer, · 1980
Earlier work this paper cites.
Improved backing-off for m-gram language modeling,
R. Kneser, H. Ney, · 1995
Earlier work this paper cites.
F. Jelinek, Statistical Methods for Speech Recognition, MIT Press, 1997
1997
Earlier work this paper cites.
Long short-term memory,
S. Hochreiter, J. Schmidhuber, · 1997
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies,
S. Hochreiter, Y. Bengio, P. Frasconi, J. Schmidhuber, · 2001
Earlier work this paper cites.
W. B. Croft, J. Lafferty, Language modeling for information retrieval, Springer Science & Business Media, 2003
2003
Earlier work this paper cites.
Quick training of probabilistic neural nets by importance sampling,
Y. Bengio, J. Senecal, · 2003
Earlier work this paper cites.
A neural probabilistic language model,
Y. Bengio, R. Ducharme, P. Vincent, C. Janvin, · 2003
Earlier work this paper cites.
Hierarchical probabilistic neural network language model,
F. Morin, Y. Bengio, · 2005
Earlier work this paper cites.
Efficient mobile phone chinese optical character recognition systems by use of heuristic fuzzy rules and bigram markov language models,
A. D. Cheok, J. Zhang, C. E. Siong, · 2008
Earlier work this paper cites.
Recurrent neural network based language model,
T. Mikolov, M. Karafiát, L. Burget, J. Cernocký, S. Khudanpur, · 2010
Earlier work this paper cites.
Tensor-train decomposition,
I. V. Oseledets, · 2011
Earlier work this paper cites.
T. Mikolov, Statistical Language Models Based on Neural Networks, Ph.D. thesis, Brno University of Technology, 2012
2012
Earlier work this paper cites.
Context dependent recurrent neural network language model,
T. Mikolov, G. Zweig, · 2012
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches,
K. Cho, B. van Merrienboer, D. Bahdanau, Y. Bengio, · 2014
Earlier work this paper cites.
Recurrent neural network regularization,
W. Zaremba, I. Sutskever, O. Vinyals, · 2014
Earlier work this paper cites.
Learning both weights and connections for efficient neural network,
S. Han, J. Pool, J. Tran, W. J. Dally, · 2015
Cited alongside, same era.
Variational dropout and the local reparameterization trick,
D. P. Kingma, T. Salimans, M. Welling, · 2015
Cited alongside, same era.
Deep fried convnets,
Z. Yang, M. Moczulski, M. Denil, N. de Freitas, A. J. Smola, L. Song, Z. Wang, · 2015
Cited alongside, same era.
Tensorizing neural networks,
A. Novikov, D. Podoprikhin, A. Osokin, D. P. Vetrov, · 2015
Cited alongside, same era.
Convolutional neural network language models,
N. Pham, G. Kruszewski, G. Boleda, · 2016
Cited alongside, same era.
Character-aware neural language models,
Y. Kim, Y. Jernite, D. Sontag, A. M. Rush, · 2016
Cited alongside, same era.
Variational dropout sparsifies deep neural networks,
D. Molchanov, A. Ashukha, D. P. Vetrov, · 2017
Later among the works it cites.
Structured bayesian pruning via log-normal multiplicative noise,
K. Neklyudov, D. Molchanov, A. Ashukha, D. P. Vetrov, · 2017
Later among the works it cites.
Bayesian sparsification of recurrent neural networks,
E. Lobacheva, N. Chirkova, D. Vetrov, · 2017
Later among the works it cites.
Compressing recurrent neural network with tensor train,
A. Tjandra, S. Sakti, S. Nakamura, · 2017
Later among the works it cites.
Long-term forecasting using tensor-train rnns,
R. Yu, S. Zheng, A. Anandkumar, Y. Yue, · 2017
Later among the works it cites.
Using the output embedding to improve language models,
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Arjovsky, A. Shah, Y. Bengio, · 2016
Cited alongside, same era.
Ultimate tensorization: compressing convolutional and FC layers alike,
T. Garipov, D. Podoprikhin, A. Novikov, D. P. Vetrov, · 2016
Cited alongside, same era.
Learning compact recurrent neural networks,
Z. Lu, V. Sindhwani, T. N. Sainath, · 2016
Cited alongside, same era.
On the compression of recurrent neural networks with an application to LVCSR acoustic modeling for embedded speech recognition,
R. Prabhavalkar, O. Alsharif, A. Bruguier, I. McGraw, · 2016
Cited alongside, same era.
Pointing the unknown words,
Ç. Gülçehre, S. Ahn, R. Nallapati, B. Zhou, Y. Bengio, · 2016
Cited alongside, same era.
Tying word vectors and word classifiers: A loss framework for language modeling,
H. Inan, K. Khosravi, R. Socher, · 2016
Cited alongside, same era.
O. Press, L. Wolf, · 2017
Later among the works it cites.
Deep neural networks performance optimization in image recognition,
A. G. Rassadin, A. V. Savchenko, · 2017
Later among the works it cites.
Recurrent highway networks,
J. G. Zilly, R. K. Srivastava, J. Koutník, J. Schmidhuber, · 2017
Later among the works it cites.
Regularizing and optimizing LSTM language models,
S. Merity, N. S. Keskar, R. Socher, · 2017
Later among the works it cites.
Breaking the softmax bottleneck: A high-rank RNN language model,
Z. Yang, Z. Dai, R. Salakhutdinov, W. W. Cohen, · 2017
Later among the works it cites.
Feature memory-based deep recurrent neural network for language modeling,
H. Deng, L. Zhang, X. Shu, · 2018
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding,
J. Devlin, M. Chang, K. Lee, K. Toutanova, · 2018
Later among the works it cites.
Online embedding compression for text classification using low rank matrix factorization,
A. Acharya, R. Goel, A. Metallinou, I. S. Dhillon, · 2018
Later among the works it cites.
Bayesian compression for natural language processing,
N. Chirkova, E. Lobacheva, D. P. Vetrov, · 2018
Later among the works it cites.
Trellis networks for sequence modeling,
S. Bai, J. Z. Kolter, V. Koltun, · 2018
Later among the works it cites.