Fetching the paper…
Reading the bibliography…
Recently, continuous cache models were proposed as extensions to recurrent neural network language models, to adapt their predictions to local changes in the data distribution.
A maximum likelihood approach to continuous speech recognition
L. R. Bahl, F. Jelinek, and R. L. Mercer · 1983
Earlier work this paper cites.
Estimation of probabilities from sparse data for the language model component of a speech recognizer
S. M. Katz · 1987
Earlier work this paper cites.
Speech recognition and the frequency of recently used words: A modified markov model for natural language
R. Kuhn · 1988
Earlier work this paper cites.
Probabilistic models of short and long distance word dependencies in running text
J. Kupiec · 1989
Earlier work this paper cites.
Finding structure in time
J. L. Elman · 1990
Earlier work this paper cites.
A cache-based natural language model for speech recognition
R. Kuhn and R. De Mori · 1990
Earlier work this paper cites.
A dynamic language model for speech recognition
F. Jelinek, B. Merialdo, S. Roukos, and M. Strauss · 1991
Earlier work this paper cites.
Adaptive language modeling using minimum discriminant estimation
S. Della Pietra, V. Della Pietra, R. L. Mercer, and S. Roukos · 1992
Earlier work this paper cites.
Variable kernel density estimation
G. R. Terrell and D. W. Scott · 1992
Earlier work this paper cites.
On the dynamic adaptation of stochastic language models
R. Kneser and V. Steinbiss · 1993
Earlier work this paper cites.
Trigger-based language models: A maximum entropy approach
R. Lau, R. Rosenfeld, and S. Roukos · 1993
Earlier work this paper cites.
Improved backing-off for m-gram language modeling
R. Kneser and H. Ney · 1995
Earlier work this paper cites.
A maximum entropy approach to adaptive statistical language modeling
R. Rosenfeld · 1996
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Multitask learning
R. Caruana · 1998
Earlier work this paper cites.
Towards better integration of semantic predictors in statistical language modeling
N. Coccaro and D. Jurafsky · 1998
Earlier work this paper cites.
Modeling long distance dependence in language: Topic mixtures versus dynamic cache models
R. M. Iyer and M. Ostendorf · 1999
Earlier work this paper cites.
Exploiting latent semantic information in statistical language modeling
J. R. Bellegarda · 2000
Earlier work this paper cites.
Maximum entropy techniques for exploiting syntactic, semantic and collocational dependencies in language modeling
S. Khudanpur and J. Wu · 2000
Earlier work this paper cites.
A bit of progress in language modeling
J. T. Goodman · 2001
Cited alongside, same era.
Similarity estimation techniques from rounding algorithms
M. S. Charikar · 2002
Cited alongside, same era.
A neural probabilistic language model
Y. Bengio, R. Ducharme, P. Vincent, and C. Jauvin · 2003
Cited alongside, same era.
Inverted files for text search engines
J. Zobel and A. Moffat · 2006
Cited alongside, same era.
Irstlm: an open source toolkit for handling large scale language models
M. Federico, N. Bertoldi, and M. Cettolo · 2008
Cited alongside, same era.
Hamming embedding and weak geometric consistency for large scale image search
H. Jegou, M. Douze, and C. Schmid · 2008
Cited alongside, same era.
From n to n+ 1: Multiclass transfer incremental learning
I. Kuzborskij, F. Orabona, and B. Caputo · 2013
Later among the works it cites.
Toward open set recognition
W. J. Scheirer, A. de Rezende Rocha, A. Sapkota, and T. E. Boult · 2013
Later among the works it cites.
N-gram counts and language models from the common crawl
C. Buck, K. Heafield, and B. van Ooyen · 2014
Later among the works it cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
J. Chung, C. Gulcehre, K. Cho, and Y. Bengio · 2014
Later among the works it cites.
Complexity of word collocation networks: A preliminary structural analysis
S. Lahiri · 2014
Later among the works it cites.
Attribute-based classification for zero-shot visual object categorization
C. H. Lampert, H. Nickisch, and S. Harmeling · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Curriculum learning
Y. Bengio, J. Louradour, R. Collobert, and J. Weston · 2009
Cited alongside, same era.
Learn ++
M. D. Muhlbaier, A. Topalis, and R. Polikar · 2009
Cited alongside, same era.
Spectral hashing
Y. Weiss, A. Torralba, and R. Fergus · 2009
Cited alongside, same era.
A theory of learning from different domains
S. Ben-David, J. Blitzer, K. Crammer, A. Kulesza, F. Pereira, and J. W. Vaughan · 2010
Cited alongside, same era.
Recurrent neural network based language model
T. Mikolov, M. Karafiát, L. Burget, J. Cernockỳ, and S. Khudanpur · 2010
Cited alongside, same era.
Iterative quantization: A procrustean approach to learning binary codes
Y. Gong and S. Lazebnik · 2011
Cited alongside, same era.
Later among the works it cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2015
Later among the works it cites.
Pointer networks
O. Vinyals, M. Fortunato, and N. Jaitly · 2015
Later among the works it cites.
Is attribute-based zero-shot learning an ill-posed strategy?
I. Alabdulmohsin, M. Cisse, and X. Zhang · 2016
Later among the works it cites.
Deep speech 2: End-to-end speech recognition in English and Mandarin
D. Amodei, R. Anubhai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, J. Chen, M. Chrzanowski, A. Coates, G. Diamos, et al · 2016
Later among the works it cites.
Fasttext.zip: Compressing text classification models
A. Joulin, E. Grave, P. Bojanowski, M. Douze, H. Jégou, and T. Mikolov · 2016
Later among the works it cites.
Exploring the limits of language modeling
R. Jozefowicz, O. Vinyals, M. Schuster, N. Shazeer, and Y. Wu · 2016
Later among the works it cites.
Building end-to-end dialogue systems using generative hierarchical neural network models
I. V. Serban, A. Sordoni, Y. Bengio, A. Courville, and J. Pineau · 2016
Later among the works it cites.
Larger-context language modelling
T. Wang and K. Cho · 2016
Later among the works it cites.
Efficient softmax approximation for GPUs
E. Grave, A. Joulin, M. Cissé, D. Grangier, and H. Jégou · 2017
Closest in time.
Improving neural language models with a continuous cache
E. Grave, A. Joulin, and N. Usunier · 2017
Closest in time.
Pointer sentinel mixture models
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2017
Closest in time.
Recurrent highway networks
J. G. Zilly, R. K. Srivastava, J. Koutník, and J. Schmidhuber · 2017
Closest in time.