Fetching the paper…
Reading the bibliography…
We introduce $k$NN-LMs, which extend a pre-trained neural language model (LM) by linearly interpolating it with a $k$-nearest neighbors ($k$NN) model.
Mbt: A memory-based part of speech tagger-generator
Walter Daelemans, Jakub Zavrel, Peter Berck, and Steven Gillis · 1996
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin · 2003
Earlier work this paper cites.
An efficient memory-based morphosyntactic tagger and parser for dutch
Antal van den Bosch, Bertjan Busser, Sander Canisius, and Walter Daelemans · 2007
Earlier work this paper cites.
Recurrent neural network based language model
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur · 2010
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou · 2017
Earlier work this paper cites.
Learning to remember rare events
Łukasz Kaiser, Ofir Nachum, Aurko Roy, and Samy Bengio · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Earlier work this paper cites.
Using the output embedding to improve language models
Ofir Press and Lior Wolf · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Lightweight adaptive mixture of neural and n-gram language models
Anton Bakhtin, Arthur Szlam, Marc’Aurelio Ranzato, and Edouard Grave · 2018
Cited alongside, same era.
Search engine guided neural machine translation
Jiatao Gu, Yong Wang, Kyunghyun Cho, and Victor OK Li · 2018
Cited alongside, same era.
Generating sentences by editing prototypes
Kelvin Guu, Tatsunori B Hashimoto, Yonatan Oren, and Percy Liang · 2018
Cited alongside, same era.
A simple cache model for image recognition
A. Emin Orhan · 2018
Cited alongside, same era.
Deep k-nearest neighbors: Towards confident, interpretable and robust deep learning
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, William W Cohen, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov · 2019
Closest in time.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Closest in time.
Dynamic evaluation of transformer language models
Ben Krause, Emmanuel Kahembwe, Iain Murray, and Steve Renals · 2019
Closest in time.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Closest in time.
Improving neural language models by segmenting, attending, and predicting the future
Hongyin Luo, Lan Jiang, Yonatan Belinkov, and James Glass · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nicolas Papernot and Patrick McDaniel · 2018
Cited alongside, same era.
Retrieve and refine: Improved sequence generation models for dialogue
Jason Weston, Emily Dinan, and Alexander H Miller · 2018
Cited alongside, same era.
Jake Zhao and Kyunghyun Cho · 2018
Cited alongside, same era.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli · 2019
Cited alongside, same era.
Unbounded cache model for online language modeling with open vocabulary
Edouard Grave, Moustapha M Cisse, and Armand Joulin
Cited in the paper.
Efficient softmax approximation for gpus
Edouard Grave, Armand Joulin, Moustapha Cissé, Hervé Jégou, et al
Cited in the paper.
Improving neural language models with a continuous cache
Edouard Grave, Armand Joulin, and Nicolas Usunier
Cited in the paper.
Closest in time.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Closest in time.
Megatron-lm: Training multi-billion parameter language models using gpu model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Closest in time.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V Le · 2019
Closest in time.
Defending against neural fake news
Rowan Zellers, Ari Holtzman, Hannah Rashkin, Ali Farhadi, Franziska Roesner, and Yejin Choi · 2019
Closest in time.