Fetching the paper…
Reading the bibliography…
This paper proposes a state-of-the-art recurrent neural network (RNN) language model that combines probability distributions computed not only from a final RNN layer but also from middle layers.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
Finding Structure in Time
Jeffrey L Elman. 1990 · 1990
Earlier work this paper cites.
A cache-based natural language model for speech recognition
Roland Kuhn and Renato De Mori. 1990 · 1990
Earlier work this paper cites.
Acceleration of Stochastic Approximation by Averaging
Boris T Polyak and Anatoli B Juditsky. 1992 · 1992
Earlier work this paper cites.
Building a Large Annotated Corpus of English: The Penn Treebank
Mitchell P. Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini. 1993 · 1993
Earlier work this paper cites.
Improved backing-off for m-gram language modeling
Reinhard Kneser and Hermann Ney. 1995 · 1995
Earlier work this paper cites.
An empirical study of smoothing techniques for language modeling
Stanley F. Chen and Joshua Goodman. 1996 · 1996
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
A maximum-entropy-inspired parser
Eugene Charniak. 2000 · 2000
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin. 2003 · 2003
Earlier work this paper cites.
Recurrent neural network based language model
Tomas Mikolov, Martin Karafiát, Lukás Burget, Jan Cernocký, and Sanjeev Khudanpur. 2010 · 2010
Earlier work this paper cites.
Annotated gigaword
Courtney Napoles, Matthew Gormley, and Benjamin Van Durme. 2012 · 2012
Earlier work this paper cites.
Distributed Representations of Words and Phrases and their Compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
Learning Word Embeddings Efficiently with Noise-Contrastive Estimation
Andriy Mnih and Koray Kavukcuoglu. 2013 · 2013
Earlier work this paper cites.
Regularization of Neural Networks using DropConnect
Li Wan, Matthew Zeiler, Sixin Zhang, Yann L Cun, and Rob Fergus. 2013 · 2013
Earlier work this paper cites.
Sequence to Sequence Learning with Neural Networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014 · 2014
Cited alongside, same era.
Recurrent neural network regularization
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals. 2014 · 2014
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
Thang Luong, Hieu Pham, and Christopher D. Manning. 2015 · 2015
Cited alongside, same era.
A Neural Attention Model for Abstractive Sentence Summarization
Alexander M. Rush, Sumit Chopra, and Jason Weston. 2015 · 2015
Cited alongside, same era.
Highway networks
Rupesh Kumar Srivastava, Klaus Greff, and Jürgen Schmidhuber. 2015 · 2015
Cited alongside, same era.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. 2015 · 2015
Cited alongside, same era.
Dynamic evaluation of neural sequence models
Ben Krause, Emmanuel Kahembwe, Iain Murray, and Steve Renals. 2017 · 2017
Later among the works it cites.
Pointer Sentinel Mixture Models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017 · 2017
Later among the works it cites.
Using the Output Embedding to Improve Language Models
Ofir Press and Lior Wolf. 2017 · 2017
Later among the works it cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, and Jeff Dean. 2017 · 2017
Later among the works it cites.
Input-to-output gate to improve rnn language models
Sho Takase, Jun Suzuki, and Masaaki Nagata. 2017 · 2017
Later among the works it cites.
Selective encoding for abstractive sentence summarization
Qingyu Zhou, Nan Yang, Furu Wei, and Ming Zhou. 2017 · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Semantically Conditioned LSTM-based Natural Language Generation for Spoken Dialogue Systems
Tsung-Hsien Wen, Milica Gasic, Nikola Mrkšić, Pei-Hao Su, David Vandyke, and Steve Young. 2015 · 2015
Cited alongside, same era.
Parsing as language modeling
Do Kook Choe and Eugene Charniak. 2016 · 2016
Cited alongside, same era.
Recurrent neural network grammars
Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, and Noah A. Smith. 2016 · 2016
Cited alongside, same era.
A Theoretically Grounded Application of Dropout in Recurrent Neural Networks
Yarin Gal and Zoubin Ghahramani. 2016 · 2016
Cited alongside, same era.
Improving neural machine translation models with monolingual data
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Improving neural parsing by disentangling model combination and reranking effects
Daniel Fried, Mitchell Stern, and Dan Klein. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
Recurrent Highway Networks
Julian Georg Zilly, Rupesh Kumar Srivastava, Jan Koutník, and Jürgen Schmidhuber. 2017 · 2017
Later among the works it cites.
Neural Architecture Search with Reinforcement Learning
Barret Zoph and Quoc V. Le. 2017 · 2017
Later among the works it cites.
Constituency parsing with a self-attentive encoder
Nikita Kitaev and Dan Klein. 2018 · 2018
Closest in time.
On the state of the art of evaluation in neural language models
Gábor Melis, Chris Dyer, and Phil Blunsom. 2018 · 2018
Closest in time.
Regularizing and Optimizing LSTM Language Models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher. 2018 · 2018
Closest in time.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Closest in time.
An empirical study of building a strong baseline for constituency parsing
Jun Suzuki, Sho Takase, Hidetaka Kamigaito, Makoto Morishita, and Masaaki Nagata. 2018 · 2018
Closest in time.
Breaking the softmax bottleneck: A high-rank RNN language model
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W. Cohen. 2018 · 2018
Closest in time.
Fraternal dropout
Konrad Zolna, Devansh Arpit, Dendi Suhubdy, and Yoshua Bengio. 2018 · 2018
Closest in time.