Fetching the paper…
Reading the bibliography…
This paper investigates the scaling properties of Recurrent Neural Network Language Models (RNNLMs).
“Complexity of exact gradient computation algorithms for recurrent neural networks,”
Ronald J. Williams, · 1989
Earlier work this paper cites.
“An application of recurrent nets to phone probability estimation,”
Anthony J. Robinson, · 1994
Earlier work this paper cites.
“Bidirectional recurrent neural networks,”
Mike Schuster and Kuldip K. Paliwal, · 1997
Earlier work this paper cites.
“Selecting articles from the language model training corpus,”
Dietrich Klakow, · 2000
Earlier work this paper cites.
BNC Consortium et al., · 2007
Earlier work this paper cites.
“Intelligent selection of language model training data,”
Robert C. Moore and William Lewis, · 2010
Earlier work this paper cites.
“Recurrent neural network based language model,”
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur, · 2010
Earlier work this paper cites.
“KenLM: Faster and smaller language model queries,”
Kenneth Heafield, · 2011
Cited alongside, same era.
“Generating text with recurrent neural networks,”
Ilya Sutskever, James Martens, and Geoffrey E Hinton, · 2011
Cited alongside, same era.
Statistical language models based on neural networks
Tomáš Mikolov, · 2012
Cited alongside, same era.
“Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude,”
T. Tieleman and G. Hinton, · 2012
Cited alongside, same era.
“TED-LIUM: an Automatic Speech Recognition dedicated corpus,”
Anthony Rousseau, Paul Deléglise, and Yannick Estève, · 2012
Cited alongside, same era.
“Building high-level features using large scale unsupervised learning,”
Quoc V Le, Rajat Monga, Matthieu Devin, Kai Chen, Greg S. Corrado, Jeff Dean, and Andrew Y. Ng, · 2013
Cited alongside, same era.
“Accelerating recurrent neural network training via two stage classes and parallelization,”
Zhiheng Huang, Geoffrey Zweig, Michael Levit, Benoit Dumoulin, Barlas Oguz, and Shawn Chang, · 2013
Later among the works it cites.
“Learning word embeddings efficiently with noise-contrastive estimation,”
Andriy Mnih and Koray Kavukcuoglu, · 2013
Later among the works it cites.
“The University of Cambridge Russian-English system at WMT13,”
J. Pino, A. Waite, T. Xiao, A. de Gispert, F. Flego, and W. Byrne, · 2013
Later among the works it cites.
“N-gram counts and language models from the common crawl,”
Christian Buck, Kenneth Heafield, and Bas van Ooyen, · 2014
Later among the works it cites.
“Efficient GPU-based training of recurrent neural network language models using spliced sentence bunch,”
Xie Chen, Yongqiang Wang, Xunying Liu, Mark JF Gales, and Philip C Woodland, · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Advances in optimizing recurrent networks,”
Yoshua Bengio, Nicolas Boulanger-Lewandowski, and Razvan Pascanu, · 2013
Cited alongside, same era.
“One Billion Word Benchmark for measuring progress in statistical language modeling,”
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson,
Cited in the paper.
“Rnnlm-recurrent neural network language modeling toolkit,”
Tomas Mikolov, Stefan Kombrink, Anoop Deoras, Lukar Burget, and J Cernocky,
Cited in the paper.
Boxun Li, Erjin Zhou, Bo Huang, Jiayi Duan, Yu Wang, Ningyi Xu, Jiaxing Zhang, and Huazhong Yang, · 2014
Later among the works it cites.