Fetching the paper…
Reading the bibliography…
Non-parametric neural language models (NLMs) learn predictive distributions of text utilizing an external datastore, which allows them to learn through explicitly memorizing the training datapoints.
A density-based algorithm for discovering clusters in large spatial databases with noise
Martin Ester, Hans-Peter Kriegel, Jörg Sander, and Xiaowei Xu. 1996 · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Similarity search in high dimensions via hashing
Aristides Gionis, Piotr Indyk, Rajeev Motwani, et al. 1999 · 1999
Earlier work this paper cites.
Video google: A text retrieval approach to object matching in videos
Josef Sivic and Andrew Zisserman. 2003 · 2003
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 2005
Earlier work this paper cites.
Fast approximate nearest neighbors with automatic algorithm configuration
Marius Muja and David G Lowe. 2009 · 2009
Earlier work this paper cites.
Product quantization for nearest neighbor search
Herve Jegou, Matthijs Douze, and Cordelia Schmid. 2010 · 2010
Earlier work this paper cites.
Recurrent neural network based language model
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur. 2010 · 2010
Earlier work this paper cites.
Iterative quantization: A procrustean approach to learning binary codes for large-scale image retrieval
Yunchao Gong, Svetlana Lazebnik, Albert Gordo, and Florent Perronnin. 2012 · 2012
Earlier work this paper cites.
Lstm neural networks for language modeling
Martin Sundermeyer, Ralf Schlüter, and Hermann Ney. 2012 · 2012
Earlier work this paper cites.
The role of hubness in clustering high-dimensional data
Nenad Tomasev, Milos Radovanovic, Dunja Mladenic, and Mirjana Ivanovic. 2013 · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Cited alongside, same era.
A neural attention model for abstractive sentence summarization
Alexander M Rush, Sumit Chopra, and Jason Weston. 2015 · 2015
Cited alongside, same era.
Learning when to skim and when to read
Alexander Johansen and Richard Socher. 2017 · 2017
Cited alongside, same era.
Six challenges for neural machine translation
Philipp Koehn and Rebecca Knowles. 2017 · 2017
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019 · 2019
Later among the works it cites.
Facebook fair’s wmt19 news translation task submission
Nathan Ng, Kyra Yee, Alexei Baevski, Myle Ott, Michael Auli, and Sergey Edunov. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
XLNet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
Later among the works it cites.
Unsupervised domain clusters in pretrained language models
Roee Aharoni and Yoav Goldberg. 2020 · 2020
Later among the works it cites.
Accelerating large-scale inference with anisotropic vector quantization
Ruiqi Guo, Philip Sun, Erik Lindgren, Quan Geng, David Simcha, Felix Chern, and Sanjiv Kumar. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generating sentences by editing prototypes
Kelvin Guu, Tatsunori B Hashimoto, Yonatan Oren, and Percy Liang. 2018 · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Character-level language modeling with deeper self-attention
Rami Al-Rfou, Dokook Choe, Noah Constant, Mandy Guo, and Llion Jones. 2019 · 2019
Cited alongside, same era.
Quicker adc: Unlocking the hidden potential of product quantization with simd
Fabien André, Anne-Marie Kermarrec, and Nicolas Le Scouarnec. 2019 · 2019
Cited alongside, same era.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli. 2019 · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Learning sparse prototypes for text generation
Junxian He, Taylor Berg-Kirkpatrick, and Graham Neubig. 2020 · 2020
Later among the works it cites.
Generalization through memorization: Nearest neighbor language models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020 · 2020
Later among the works it cites.
Neurips 2020 efficientqa competition: Systems, analyses and lessons learned
Sewon Min, Jordan Boyd-Graber, Chris Alberti, Danqi Chen, Eunsol Choi, Michael Collins, Kelvin Guu, Hannaneh Hajishirzi, Kenton Lee, Jennimaria Palomaki, et al. 2020 · 2020
Later among the works it cites.
Nearest neighbor machine translation
Urvashi Khandelwal, Angela Fan, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2021 · 2021
Closest in time.
Adaptive semiparametric language models
Dani Yogatama, Cyprien de Masson d’Autume, and Lingpeng Kong. 2021 · 2021
Closest in time.