Fetching the paper…
Reading the bibliography…
We propose an extension to neural network language models to adapt their prediction to the recent history.
A maximum likelihood approach to continuous speech recognition
Lalit R Bahl, Frederick Jelinek, and Robert L Mercer · 1983
Earlier work this paper cites.
Estimation of probabilities from sparse data for the language model component of a speech recognizer
Slava M Katz · 1987
Earlier work this paper cites.
Speech recognition and the frequency of recently used words: A modified markov model for natural language
Roland Kuhn · 1988
Earlier work this paper cites.
Probabilistic models of short and long distance word dependencies in running text
Julien Kupiec · 1989
Earlier work this paper cites.
Finding structure in time
Jeffrey L Elman · 1990
Earlier work this paper cites.
A cache-based natural language model for speech recognition
Roland Kuhn and Renato De Mori · 1990
Earlier work this paper cites.
Backpropagation through time: what it does and how to do it
Paul J Werbos · 1990
Earlier work this paper cites.
An efficient gradient-based algorithm for on-line training of recurrent network trajectories
Ronald J Williams and Jing Peng · 1990
Earlier work this paper cites.
A dynamic language model for speech recognition
Frederick Jelinek, Bernard Merialdo, Salim Roukos, and Martin Strauss · 1991
Earlier work this paper cites.
Adaptive language modeling using minimum discriminant estimation
Stephen Della Pietra, Vincent Della Pietra, Robert L Mercer, and Salim Roukos · 1992
Earlier work this paper cites.
The mathematics of statistical machine translation: Parameter estimation
Peter F Brown, Vincent J Della Pietra, Stephen A Della Pietra, and Robert L Mercer · 1993
Earlier work this paper cites.
On the dynamic adaptation of stochastic language models
Reinhard Kneser and Volker Steinbiss · 1993
Earlier work this paper cites.
Trigger-based language models: A maximum entropy approach
Raymond Lau, Ronald Rosenfeld, and Salim Roukos · 1993
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell P Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 1993
Earlier work this paper cites.
Improved backing-off for m-gram language modeling
Reinhard Kneser and Hermann Ney · 1995
Earlier work this paper cites.
A maximum entropy approach to adaptive statistical language modeling
Ronald Rosenfeld · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Towards better integration of semantic predictors in statistical language modeling
Noah Coccaro and Daniel Jurafsky · 1998
Cited alongside, same era.
Modeling long distance dependence in language: Topic mixtures versus dynamic cache models
Rukmini M Iyer and Mari Ostendorf · 1999
Cited alongside, same era.
Exploiting latent semantic information in statistical language modeling
Jerome R Bellegarda · 2000
Cited alongside, same era.
Maximum entropy techniques for exploiting syntactic, semantic and collocational dependencies in language modeling
Sanjeev Khudanpur and Jun Wu · 2000
Cited alongside, same era.
Dialogue act modeling for automatic tagging and recognition of conversational speech
Andreas Stolcke, Noah Coccaro, Rebecca Bates, Paul Taylor, Carol Van Ess-Dykema, Klaus Ries, Elizabeth Shriberg, Daniel Jurafsky, Rachel Martin, and Marie Meteer · 2000
Cited alongside, same era.
Evaluating prerequisite qualities for learning end-to-end dialog systems
Jesse Dodge, Andreea Gane, Xiang Zhang, Antoine Bordes, Sumit Chopra, Alexander Miller, Arthur Szlam, and Jason Weston · 2015
Later among the works it cites.
A theoretically grounded application of dropout in recurrent neural networks
Yarin Gal and Zoubin Ghahramani · 2015
Later among the works it cites.
Learning to transduce with unbounded memory
Edward Grefenstette, Karl Moritz Hermann, Mustafa Suleyman, and Phil Blunsom · 2015
Later among the works it cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom · 2015
Later among the works it cites.
Inferring algorithmic patterns with stack-augmented recurrent nets
Armand Joulin and Tomas Mikolov · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Joshua T Goodman · 2001
Cited alongside, same era.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin · 2003
Cited alongside, same era.
Recurrent neural network based language model
Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernockỳ, and Sanjeev Khudanpur · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Cited alongside, same era.
Empirical evaluation and combination of advanced language modeling techniques
Tomas Mikolov, Anoop Deoras, Stefan Kombrink, Lukas Burget, and Jan Cernockỳ · 2011
Cited alongside, same era.
Context dependent recurrent neural network language model
Tomas Mikolov and Geoffrey Zweig · 2012
Cited alongside, same era.
Speech recognition with deep recurrent neural networks
A. Graves, A. Mohamed, and G. Hinton · 2013
Cited alongside, same era.
End-to-end memory networks
Sainbayar Sukhbaatar, Szlam Arthur, Jason Weston, and Rob Fergus · 2015
Later among the works it cites.
Pointer networks
Oriol Vinyals, Meire Fortunato, and Navdeep Jaitly · 2015
Later among the works it cites.
Larger-context language modelling
Tian Wang and Kyunghyun Cho · 2015
Later among the works it cites.
A thorough examination of the cnn/daily mail reading comprehension task
Danqi Chen, Jason Bolton, and Christopher D Manning · 2016
Closest in time.
Efficient softmax approximation for gpus
Edouard Grave, Armand Joulin, Moustapha Cissé, David Grangier, and Hervé Jégou · 2016
Closest in time.
Caglar Gulcehre, Sungjin Ahn, Ramesh Nallapati, Bowen Zhou, and Yoshua Bengio · 2016
Closest in time.
Exploring the limits of language modeling
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu · 2016
Closest in time.
Text understanding with the attention sum reader network
Rudolf Kadlec, Martin Schmid, Ondrej Bajgar, and Jan Kleindienst · 2016
Closest in time.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Closest in time.
The lambada dataset: Word prediction requiring a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández · 2016
Closest in time.
Julian Georg Zilly, Rupesh Kumar Srivastava, Jan Koutník, and Jürgen Schmidhuber · 2016
Closest in time.