Fetching the paper…
Reading the bibliography…
Fixed-vocabulary language models fail to account for one of the most characteristic statistical facts of natural language: the frequent creation and reuse of new word types.
Information retrieval: Computational and theoretical aspects
Harold Stanley Heaps. 1978 · 1978
Earlier work this paper cites.
A cache-based natural language model for speech recognition
Roland Kuhn and Renato De Mori. 1990 · 1990
Earlier work this paper cites.
Poisson mixtures
Kenneth W Church and William A Gale. 1995 · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Empirical estimates of adaptation: the chance of two Noriegas is closer to p / 2 p/2 than p 2 p^{2}
Kenneth W Church. 2000 · 2000
Earlier work this paper cites.
A hierarchical Bayesian language model based on Pitman-Yor processes
Yee Whye Teh. 2006 · 2006
Earlier work this paper cites.
A Bayesian framework for word segmentation: Exploring the effects of context
Sharon Goldwater, Thomas L Griffiths, and Mark Johnson. 2009 · 2009
Earlier work this paper cites.
Recurrent neural network based language model
Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernockỳ, and Sanjeev Khudanpur. 2010 · 2010
Earlier work this paper cites.
Generating text with recurrent neural networks
Ilya Sutskever, James Martens, and Geoffrey E Hinton. 2011 · 2011
Earlier work this paper cites.
Subword language modeling with neural networks
Tomáš Mikolov, Ilya Sutskever, Anoop Deoras, Hai-Son Le, Stefan Kombrink, and Jan Cernocky. 2012 · 2012
Cited alongside, same era.
Knowledge-rich morphological priors for bayesian language models
Victor Chahuneau, Noah A. Smith, and Chris Dyer. 2013 · 2013
Cited alongside, same era.
Generating sequences with recurrent neural networks
Alex Graves. 2013 · 2013
Cited alongside, same era.
A clockwork RNN
Jan Koutnik, Klaus Greff, Faustino Gomez, and Juergen Schmidhuber. 2014 · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba. 2015 · 2015
On multiplicative integration with recurrent neural networks
Yuhuai Wu, Saizheng Zhang, Ying Zhang, Yoshua Bengio, and Ruslan R Salakhutdinov. 2016 · 2016
Later among the works it cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le. 2016 · 2016
Later among the works it cites.
Hierarchical multiscale recurrent neural networks
Junyoung Chung, Sungjin Ahn, and Yoshua Bengio. 2017 · 2017
Closest in time.
Recurrent batch normalization
Tim Cooijmans, Nicolas Ballas, César Laurent, Çağlar Gülçehre, and Aaron Courville. 2017 · 2017
Closest in time.
Improving neural language models with a continuous cache
Edouard Grave, Armand Joulin, and Nicolas Usunier. 2017 · 2017
Closest in time.
Hypernetworks
David Ha, Andrew Dai, and Quoc V Le. 2017 · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Finding function in form: Compositional character models for open vocabulary word representation
Wang Ling, Tiago Luís, Luís Marujo, Ramón Fernandez Astudillo, Silvio Amir, Chris Dyer, Alan W Black, and Isabel Trancoso. 2015 · 2015
Cited alongside, same era.
A hierarchical recurrent encoder-decoder for generative context-aware query suggestion
Alessandro Sordoni, Yoshua Bengio, Hossein Vahabi, Christina Lioma, Jakob Grue Simonsen, and Jian-Yun Nie. 2015 · 2015
Cited alongside, same era.
Recurrent dropout without memory loss
Stanislau Semeniuta, Aliaksei Severyn, and Erhardt Barth. 2016 · 2016
Cited alongside, same era.
Zoneout: Regularizing rnns by randomly preserving hidden activations
David Krueger, Tegan Maharaj, János Kramár, Mohammad Pezeshki, Nicolas Ballas, Nan Rosemary Ke, Anirudh Goyal, Yoshua Bengio, Hugo Larochelle, Aaron Courville, et al. 2017 · 2017
Closest in time.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017 · 2017
Closest in time.