An inequality with application to statistical estimation for probabilistic functions of markov processes and to a model for ecology
LE Baum and JA Eagon · 1967
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell P Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 1993
Earlier work this paper cites.
First-order versus second-order single-layer recurrent neural networks
Mark W Goudreau, C Lee Giles, Srimat T Chakradhar, and D Chen · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Refining hidden markov models with recurrent neural networks
T. Wessels and C. W. Omlin · 2000
Earlier work this paper cites.
Weighted finite-state transducers in speech recognition
Mehryar Mohri, Fernando Pereira, and Michael Riley · 2002
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
Generating text with recurrent neural networks
Ilya Sutskever, James Martens, and Geoffrey E Hinton · 2011
Earlier work this paper cites.
Subword language modeling with neural networks
Tomáš Mikolov, Ilya Sutskever, Anoop Deoras, Hai-Son Le, and Stefan Kombrink · 2012
Earlier work this paper cites.
Regularization and nonlinearities for neural language models: when are they needed?
Original
Marius Pachitariu and Maneesh Sahani · 2013
Earlier work this paper cites.