Fetching the paper…
Reading the bibliography…
Semiparametric language models (LMs) have shown promise in various Natural Language Processing (NLP) tasks.
Augmenting Self-attention with Persistent Memory, 2019
Sainbayar Sukhbaatar, Edouard Grave, Guillaume Lample, Herve Jegou, and Armand Joulin · 1907
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
Robert M. French · 1999
Earlier work this paper cites.
REALM: Retrieval-Augmented Language Model Pre-Training, 2020
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang · 2002
Earlier work this paper cites.
Analyzing the forgetting problem in the pretrain-finetuning of dialogue response models
Tianxing He, Jun Liu, Kyunghyun Cho, Myle Ott, Bing Liu, James Glass, and Fuchun Peng · 2019
Earlier work this paper cites.
Billion-scale similarity search with GPUs
Jeff Johnson, Matthijs Douze, and Hervé Jégou · 2019
Earlier work this paper cites.
Continual learning for sentence representations using conceptors
Tianlin Liu, Lyle Ungar, and João Sedoc · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Recall and learn: Fine-tuning deep pretrained language models with less forgetting
Sanyuan Chen, Yutai Hou, Yiming Cui, Wanxiang Che, Ting Liu, and Xiangzhan Yu · 2020
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis · 2020
Earlier work this paper cites.
S2ORC: The semantic scholar open research corpus
Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel Weld · 2020
Earlier work this paper cites.
LAMOL: language modeling for lifelong language learning
Fan-Keng Sun, Cheng-Hao Ho, and Hung-Yi Lee · 2020
Earlier work this paper cites.
Efficient nearest neighbor language models
Junxian He, Graham Neubig, and Taylor Berg-Kirkpatrick · 2021
Cited alongside, same era.
Learning kernel-smoothed machine translation with retrieved examples
Qingnan Jiang, Mingxuan Wang, Jun Cao, Shanbo Cheng, Shujian Huang, and Lei Li · 2021
Cited alongside, same era.
Learn continually, generalize rapidly: Lifelong knowledge accumulation for few-shot learning
Xisen Jin, Bill Yuchen Lin, Mohammad Rostami, and Xiang Ren · 2021
Cited alongside, same era.
Nearest neighbor machine translation
Urvashi Khandelwal, Angela Fan, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis · 2021
Cited alongside, same era.
Adaptive semiparametric language models
Dani Yogatama, Cyprien de Masson d’Autume, and Lingpeng Kong · 2021
Cited alongside, same era.
Towards continual knowledge learning of language models
Joel Jang, Seonghyeon Ye, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun Kim, Stanley Jungkyu Choi, and Minjoon Seo · 2022
Later among the works it cites.
Lifelong pretraining: Continually adapting language models to emerging corpora
Xisen Jin, Dejiao Zhang, Henghui Zhu, Wei Xiao, Shang-Wen Li, Xiaokai Wei, Andrew Arnold, and Xiang Ren · 2022
Later among the works it cites.
RealTime QA: What’s the Answer Right Now?, 2022
Jungo Kasai, Keisuke Sakaguchi, Yoichi Takahashi, Ronan Le Bras, Akari Asai, Xinyan Yu, Dragomir Radev, Noah A. Smith, Yejin Choi, and Kentaro Inui · 2022
Later among the works it cites.
Continual Training of Language Models for Few-Shot Learning, 2022
Zixuan Ke, Haowei Lin, Yijia Shao, Hu Xu, Lei Shu, and Bing Liu · 2022
Later among the works it cites.
LFPT5: A unified framework for lifelong few-shot language learning based on prompt tuning of T5
Chengwei Qin and Shafiq Joty · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xin Zheng, Zhirui Zhang, Junliang Guo, Shujian Huang, Boxing Chen, Weihua Luo, and Jiajun Chen · 2021
Cited alongside, same era.
Neuro-symbolic language modeling with automaton-augmented retrieval
Uri Alon, Frank F. Xu, Junxian He, Sudipta Sengupta, Dan Roth, and Graham Neubig · 2022
Cited alongside, same era.
Adaptation Approaches for Nearest Neighbor Language Models, 2022
Rishabh Bhardwaj, George Polovets, and Monica Sunkara · 2022
Cited alongside, same era.
Improving language models by retrieving from trillions of tokens
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Chris Jones, Albin Cassirer, Andy Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack W. Rae, Erich Elsen, and Laurent Sifre · 2022
Cited alongside, same era.
You can’t pick your neighbors, or can you? When and how to rely on retrieval in the $k$NN-LM, 2022
Andrew Drozdov, Shufan Wang, Razieh Rahimi, Andrew McCallum, Hamed Zamani, and Mohit Iyyer · 2022
Cited alongside, same era.
Plug and play knowledge distillation for knn-lm with external logits
Xuyang Jin, Tao Ge, and Furu Wei
Cited in the paper.
Nearest neighbor zero-shot inference
Weijia Shi, Julian Michael, Suchin Gururangan, and Luke Zettlemoyer
Cited in the paper.
Later among the works it cites.
ELLE: Efficient lifelong pre-training for emerging data
Yujia Qin, Jiajie Zhang, Yankai Lin, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou · 2022
Later among the works it cites.
Continual-t0: Progressively instructing 50+ tasks to language models without forgetting
Thomas Scialom, Tuhin Chakrabarty, and Smaranda Muresan · 2022
Later among the works it cites.
Nearest Neighbor Language Models for Stylistic Controllable Generation, 2022
Severino Trotta, Lucie Flek, and Charles Welch · 2022
Later among the works it cites.
Memorizing transformers
Yuhuai Wu, Markus Norman Rabe, DeLesley Hutchins, and Christian Szegedy · 2022
Later among the works it cites.
ConTinTin: Continual learning from task instructions
Wenpeng Yin, Jia Li, and Caiming Xiong · 2022
Later among the works it cites.