Fetching the paper…
Reading the bibliography…
We introduce adaptive input representations for neural language modeling which extend the adaptive softmax of Grave et al.
Classes for Fast Maximum Entropy Training
Joshua Goodman · 2001
Earlier work this paper cites.
A Neural Probabilistic Language Model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin · 2003
Earlier work this paper cites.
Hierarchical Probabilistic Neural Network Language Model
Frederic Morin and Yoshua Bengio · 2005
Earlier work this paper cites.
Recurrent Neural Network based Language Model
Tomáš Mikolov, Karafiát Martin, Lukáš Burget, Jan Cernocký, and Sanjeev Khudanpur · 2010
Earlier work this paper cites.
Extensions of Recurrent Neural Network Language Model
Tomáš Mikolov, Stefan Kombrink, Lukáš Burget, Jan Cernocký, and Sanjeev Khudanpur · 2011
Earlier work this paper cites.
Deep Neural Network Language Models
Ebru Arisoy, Tara N. Sainath, Brian Kingsbury, and Bhuvana Ramabhadran · 2012
Earlier work this paper cites.
Large, Pruned or Continuous Space Language Models on a GPU for Statistical Machine Translation
Holger Schwenk, Anthony Rousseau, and Mohammed Attik · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George E. Dahl, and Geoffrey E. Hinton · 2013
Earlier work this paper cites.
Decoding with Large-scale Neural Language Models improves Translation
Ashish Vaswani, Yinggong Zhao, Victoria Fossum, and David Chiang · 2013
Earlier work this paper cites.
Pragmatic neural language modelling in machine translation
Paul Baltescu and Phil Blunsom · 2015
Cited alongside, same era.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Character-aware neural language models
Yoon Kim, Yacine Jernite, David Sontag, and Alexander M. Rush · 2015
Cited alongside, same era.
Strategies for training large vocabulary neural language models
Wenlin Chen, David Grangier, and Michael Auli · 2016
Cited alongside, same era.
Improving neural language models with a continuous cache
Edouard Grave, Armand Joulin, and Nicolas Usunier · 2016
Cited alongside, same era.
Sparse non-negative matrix language modeling
Noam Shazeer, Joris Pelemans, and Ciprian Chelba · 2016
Later among the works it cites.
Language modeling with gated convolutional networks
Yann N. Dauphin, Angela Fan, Michael Auli, and David Grangier · 2017
Later among the works it cites.
Efficient softmax approximation for gpus
Edouard Grave, Armand Joulin, Moustapha Cissé, David Grangier, and Hervé Jégou · 2017
Later among the works it cites.
Using the output embedding to improve language models
Ofir Press and Lior Wolf · 2017
Later among the works it cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, and Jeff Dean · 2017
Later among the works it cites.
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hakan Inan, Khashayar Khosravi, and Richard Socher · 2016
Cited alongside, same era.
Exploring the limits of language modeling
Rafal Józefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu · 2016
Cited alongside, same era.
Character-aware neural language models
Yoon Kim, Yacine Jernite, David Sontag, and Alexander M Rush · 2016
Cited alongside, same era.
SGDR: stochastic gradient descent with restarts
Ilya Loshchilov and Frank Hutter · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Cited alongside, same era.
Analyzing uncertainty in neural machine translation
Myle Ott, Michael Auli, David Grangier, and Marc’Aurelio Ranzato
Cited in the paper.
Later among the works it cites.
Character-level language modeling with deeper self-attention
Rami Al-Rfou, Dokook Choe, Noah Constant, Mandy Guo, and Llion Jones · 2018
Closest in time.
Neural lattice language models
Jacob Buckman and Graham Neubig · 2018
Closest in time.
An analysis of neural language modeling at multiple scales
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2018
Closest in time.
Spell once, summon anywhere: A two-level open-vocabulary language model
Sebastian J. Mielke and Jason Eisner · 2018
Closest in time.
Fast parametric learning with activation memorization
Jack W. Rae, Chris Dyer, Peter Dayan, and Timothy P. Lillicrap · 2018
Closest in time.