Fetching the paper…
Reading the bibliography…
Many recent efforts augment language models with retrieval, by adding retrieved data to the input context.
Consistent nonparametric regression
Charles J Stone · 1977
Earlier work this paper cites.
Robust locally weighted regression and smoothing scatterplots
William S Cleveland · 1979
Earlier work this paper cites.
Locally weighted regression: an approach to regression analysis by local fitting
William S Cleveland and Susan J Devlin · 1988
Earlier work this paper cites.
Local learning algorithms
Léon Bottou and Vladimir Vapnik · 1992
Earlier work this paper cites.
An effective bandwidth selector for local least squares regression
David Ruppert, Simon J Sheather, and Matthew P Wand · 1995
Earlier work this paper cites.
Locally weighted learning
Christopher G Atkeson, Andrew W Moore, and Stefan Schaal · 1997
Earlier work this paper cites.
Learning by transduction
Alexander Gammerman, Volodya Vovk, and Vladimir Vapnik · 1998
Earlier work this paper cites.
Learning to classify text using support vector machines , volume 668
Thorsten Joachims · 2002
Earlier work this paper cites.
Large scale transductive svms
Ronan Collobert, Fabian Sinz, Jason Weston, Léon Bottou, and Thorsten Joachims · 2006
Earlier work this paper cites.
Svm-knn: Discriminative nearest neighbor classification for visual category recognition
Hao Zhang, Alexander C Berg, Michael Maire, and Jitendra Malik · 2006
Earlier work this paper cites.
The nature of statistical learning theory
Vladimir Vapnik · 2013
Cited alongside, same era.
Dynamic evaluation of neural sequence models
Ben Krause, Emmanuel Kahembwe, Iain Murray, and Steve Renals · 2018
Cited alongside, same era.
Billion-scale similarity search with GPUs
Jeff Johnson, Matthijs Douze, and Hervé Jégou · 2019
Cited alongside, same era.
Generalization through memorization: Nearest neighbor language models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis · 2019
Cited alongside, same era.
Dynamic evaluation of transformer language models
Ben Krause, Emmanuel Kahembwe, Iain Murray, and Steve Renals · 2019
Cited alongside, same era.
Test-time training with self-supervision for generalization under distribution shifts
Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei Efros, and Moritz Hardt · 2020
Later among the works it cites.
A framework for few-shot language model evaluation
Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou · 2021
Later among the works it cites.
Demix layers: Disentangling domains for modular language modeling
Suchin Gururangan, Mike Lewis, Ari Holtzman, Noah A Smith, and Luke Zettlemoyer · 2021
Later among the works it cites.
Efficient nearest neighbor language models
Junxian He, Graham Neubig, and Taylor Berg-Kirkpatrick · 2021
Later among the works it cites.
Neuro-symbolic language modeling with automaton-augmented retrieval
Uri Alon, Frank Xu, Junxian He, Sudipta Sengupta, Dan Roth, and Graham Neubig · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Cited alongside, same era.
Retrieval augmented language model pre-training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang · 2020
Cited alongside, same era.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Cited alongside, same era.
Test-time training on video streams
Renhao Wang, Yu Sun, Yossi Gandelsman, Xinlei Chen, Alexei A Efros, and Xiaolong Wang
Cited in the paper.
Shuohang Wang, Yichong Xu, Yuwei Fang, Yang Liu, Siqi Sun, Ruochen Xu, Chenguang Zhu, and Michael Zeng
Cited in the paper.
Later among the works it cites.
Improving language models by retrieving from trillions of tokens
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al · 2022
Later among the works it cites.
Test-time training with masked autoencoders
Yossi Gandelsman, Yu Sun, Xinlei Chen, and Alexei A. Efros · 2022
Later among the works it cites.
Few-shot learning with retrieval augmented language models
Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave · 2022
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Closest in time.
Learning to (learn at test time)
Yu Sun, Xinhao Li, Karan Dalal, Chloe Hsu, Sanmi Koyejo, Carlos Guestrin, Xiaolong Wang, Tatsunori Hashimoto, and Xinlei Chen · 2023
Closest in time.