Fetching the paper…
Reading the bibliography…
Recurrent neural network (RNN) based character-level language models (CLMs) are extremely useful for modeling out-of-vocabulary words by nature.
“A method of solving a convex programming problem with convergence rate O (1/k2),”
Yurii Nesterov, · 1983
Earlier work this paper cites.
“A statistical approach to machine translation,”
Peter F Brown, John Cocke, Stephen A Della Pietra, Vincent J Della Pietra, Fredrick Jelinek, John D Lafferty, Robert L Mercer, and Paul S Roossin, · 1990
Earlier work this paper cites.
“Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition,”
John S Bridle, · 1990
Earlier work this paper cites.
“Finding structure in time,”
Jeffrey L Elman, · 1990
Earlier work this paper cites.
“Backpropagation through time: what it does and how to do it,”
Paul J Werbos, · 1990
Earlier work this paper cites.
“An efficient gradient-based algorithm for on-line training of recurrent network trajectories,”
Ronald J Williams and Jing Peng, · 1990
Earlier work this paper cites.
“The design for the Wall Street Journal-based CSR corpus,”
Douglas B Paul and Janet M Baker, · 1992
Earlier work this paper cites.
Fundamentals of speech recognition
Lawrence Rabiner and Biing-Hwang Juang, · 1993
Earlier work this paper cites.
“Building a large annotated corpus of english: The penn treebank,”
Mitchell P Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini, · 1993
Earlier work this paper cites.
“Improved backing-off for m-gram language modeling,”
Reinhard Kneser and Hermann Ney, · 1995
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Gated word-character recurrent language model,”
Yasumasa Miyamoto and Kyunghyun Cho, · 1997
Cited alongside, same era.
“Learning to forget: Continual prediction with LSTM,”
Felix A Gers, Jürgen Schmidhuber, and Fred Cummins, · 2000
Cited alongside, same era.
“Learning precise timing with LSTM recurrent networks,”
Felix A Gers, Nicol N Schraudolph, and Jürgen Schmidhuber, · 2003
Cited alongside, same era.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Cited alongside, same era.
“Recurrent neural network based language model.,”
Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernockỳ, and Sanjeev Khudanpur, · 2010
Cited alongside, same era.
“Generating text with recurrent neural networks,”
Ilya Sutskever, James Martens, and Geoffrey E Hinton, · 2011
“One billion word benchmark for measuring progress in statistical language modeling,”
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson, · 2013
Later among the works it cites.
“Finding function in form: Compositional character models for open vocabulary word representation,”
Wang Ling, Tiago Luís, Luís Marujo, Ramón Fernandez Astudillo, Silvio Amir, Chris Dyer, Alan W Black, and Isabel Trancoso, · 2015
Later among the works it cites.
“Single stream parallelization of generalized LSTM-like RNNs on a GPU,”
Kyuyeon Hwang and Wonyong Sung, · 2015
Later among the works it cites.
“Sparse non-negative matrix language modeling for skip-grams,”
Noam Shazeer, Joris Pelemans, and Ciprian Chelba, · 2015
Later among the works it cites.
“Exploring the limits of language modeling,”
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu, · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“ADADELTA: An adaptive learning rate method,”
Matthew D Zeiler, · 2012
Cited alongside, same era.
“Improving neural networks by preventing co-adaptation of feature detectors,”
Geoffrey E Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan R Salakhutdinov, · 2012
Cited alongside, same era.
Statistical language models based on neural networks
Tomáš Mikolov, · 2012
Cited alongside, same era.
“Context dependent recurrent neural network language model,”
Tomas Mikolov and Geoffrey Zweig, · 2012
Cited alongside, same era.
“Training and analysing deep recurrent neural networks,”
Michiel Hermans and Benjamin Schrauwen, · 2013
Cited alongside, same era.
Closest in time.
“Character-level incremental speech recognition with recurrent neural networks,”
Kyuyeon Hwang and Wonyong Sung, · 2016
Closest in time.
“Character-aware neural language models,”
Yoon Kim, Yacine Jernite, David Sontag, and Alexander M Rush, · 2016
Closest in time.
“Character-based neural machine translation,”
Wang Ling, Isabel Trancoso, Chris Dyer, and Alan W Black, · 2016
Closest in time.
“BlackOut: Speeding up recurrent neural network language models with very large vocabularies,”
Shihao Ji, SVN Vishwanathan, Nadathur Satish, Michael J Anderson, and Pradeep Dubey, · 2016
Closest in time.
“LightRNN: Memory and computation-efficient recurrent neural networks,”
Xiang Li, Tao Qin, Jian Yang, Xiaolin Hu, and Tieyan Liu, · 2016
Closest in time.