Fetching the paper…
Reading the bibliography…
We know very little about how neural language models (LM) use prior linguistic context.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell P Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini. 1993 · 2004
Earlier work this paper cites.
Syntactic topic models
Jordan Boyd-Graber and David Blei. 2009 · 2009
Earlier work this paper cites.
Recurrent neural network based language model
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur. 2010 · 2010
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Alex Graves. 2013 · 2013
Earlier work this paper cites.
Regularization of neural networks using dropconnect
Li Wan, Matthew Zeiler, Sixin Zhang, Yann Le Cun, and Rob Fergus. 2013 · 2013
Earlier work this paper cites.
The stanford corenlp natural language processing toolkit
Christopher Manning, Mihai Surdeanu, John Bauer, Jenny Finkel, Steven Bethard, and David McClosky. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
A theoretically grounded application of dropout in recurrent neural networks
Yarin Gal and Zoubin Ghahramani. 2016 · 2016
Earlier work this paper cites.
Contextual lstm (clstm) models for large scale nlp tasks
Shalini Ghosh, Oriol Vinyals, Brian Strope, Scott Roy, Tom Dean, and Larry Heck. 2016 · 2016
Cited alongside, same era.
The goldilocks principle: Reading children’s books with explicit memory representations
Felix Hill, Antoine Bordes, Sumit Chopra, and Jason Weston. 2016 · 2016
Cited alongside, same era.
Exploring the limits of language modeling
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. 2016 · 2016
Cited alongside, same era.
Visualizing and understanding neural models in nlp
Jiwei Li, Xinlei Chen, Eduard Hovy, and Dan Jurafsky. 2016 · 2016
Cited alongside, same era.
Assessing the ability of lstms to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016 · 2016
Language modeling with gated convolutional networks
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier. 2017 · 2017
Later among the works it cites.
Tying word vectors and word classifiers: A loss framework for language modeling
Hakan Inan, Khashayar Khosravi, and Richard Socher. 2017 · 2017
Later among the works it cites.
Topically Driven Neural Language Model
Jey Han Lau, Timothy Baldwin, and Trevor Cohn. 2017 · 2017
Later among the works it cites.
Pointer Sentinel Mixture Models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017 · 2017
Later among the works it cites.
On the State of the Art of Evaluation in Neural Language Models
Gabor Melis, Chris Dyer, and Phil Blunsom. 2018 · 2018
Closest in time.
Regularizing and Optimizing LSTM Language Models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Larger-Context Language Modelling with Recurrent Neural Network
Tian Wang and Kyunghyun Cho. 2016 · 2016
Cited alongside, same era.
Fine-grained analysis of sentence embeddings using auxiliary prediction tasks
Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg. 2017 · 2017
Cited alongside, same era.
N-gram language modeling using recurrent neural network estimation
Ciprian Chelba, Mohammad Norouzi, and Samy Bengio. 2017 · 2017
Cited alongside, same era.
Unbounded cache model for online language modeling with open vocabulary
Edouard Grave, Moustapha M Cisse, and Armand Joulin. 2017a
Cited in the paper.
Improving Neural Language Models with a Continuous Cache
Edouard Grave, Armand Joulin, and Nicolas Usunier. 2017b
Cited in the paper.
Closest in time.
Breaking the softmax bottleneck: a high-rank rnn language model
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W Cohen. 2018 · 2018
Closest in time.
Using the output embedding to improve language models
Ofir Press and Lior Wolf. 2017 · 2025
Closest in time.