Fetching the paper…
Reading the bibliography…
Keeping large language models factually up-to-date is crucial for deployment, yet costly retraining remains a challenge.
Eli5: Long form question answering, 2019
Fan, A., Jernite, Y., Perez, E., Grangier, D., Weston, J., and Auli, M · 1907
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Reimers, N. and Gurevych, I · 1908
Earlier work this paper cites.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Realm: Retrieval-augmented language model pre-training, 2020
Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M.-W · 2002
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering, 2020
Karpukhin, V., Oğuz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., and tau Yih, W · 2004
Earlier work this paper cites.
Is multihop qa in dire condition? measuring and reducing disconnected reasoning, 2020
Trivedi, H., Balasubramanian, N., Khot, T., and Sabharwal, A · 2005
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension, 2017
Joshi, M., Choi, E., Weld, D. S., and Zettlemoyer, L · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., Hassabis, D., Clopath, C., Kumaran, D., and Hadsell, R · 2017
Earlier work this paper cites.
Zero-shot relation extraction via reading comprehension
Levy, O., Seo, M., Choi, E., and Zettlemoyer, L · 2017
Earlier work this paper cites.
Continual learning through synaptic intelligence, 2017
Zenke, F., Poole, B., and Ganguli, S · 2017
Earlier work this paper cites.
Ms marco: A human generated machine reading comprehension dataset, 2018
Bajaj, P., Campos, D., Craswell, N., Deng, L., Gao, J., Liu, X., Majumder, R., McNamara, A., Mitra, B., Nguyen, T., Rosenberg, M., Song, X., Stoica, A., Tiwary, S., and Wang, T · 2018
Earlier work this paper cites.
Annoy: Approximate Nearest Neighbors in C++/Python , 2018
Bernhardsson, E · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering, 2018
Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W. W., Salakhutdinov, R., and Manning, C. D · 2018
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Kelcey, M., Devlin, J., Lee, K., Toutanova, K. N., Jones, L., Chang, M.-W., Dai, A., Uszkoreit, J., Le, Q., and Petrov, S · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., and Kiela, D · 2020
Earlier work this paper cites.
Mpnet: Masked and permuted pre-training for language understanding
Song, K., Tan, X., Qin, T., Lu, J., and Liu, T.-Y · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models, 2021
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Fast model editing at scale
Mitchell, E., Lin, C., Bosselut, A., Finn, C., and Manning, C. D · 2021
Earlier work this paper cites.
ArchivalQA: A large-scale benchmark dataset for open domain question answering over historical news collections
Wang, J., Jatowt, A., and Yoshikawa, M · 2021
Earlier work this paper cites.
Improving language models by retrieving from trillions of tokens, 2022
Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., van den Driessche, G., Lespiau, J.-B., Damoc, B., Clark, A., de Las Casas, D., Guy, A., Menick, J., Ring, R., Hennigan, T., Huang, S., Maggiore, L., Jones, C., Cassirer, A., Brock, A., Paganini, M., Irving, G., Vinyals, O., Osindero, S., Simonyan, K., Rae, J. W., Elsen, E., and Sifre, L · 2022
Earlier work this paper cites.
True: Re-evaluating factual consistency evaluation, 2022
Honovich, O., Aharoni, R., Herzig, J., Taitelbaum, H., Kukliansy, D., Cohen, V., Scialom, T., Szpektor, I., Hassidim, A., and Matias, Y · 2022
Earlier work this paper cites.
Patching open-vocabulary models by interpolating weights, 2022
Ilharco, G., Wortsman, M., Gadre, S. Y., Song, S., Hajishirzi, H., Kornblith, S., Farhadi, A., and Schmidt, L · 2022
Earlier work this paper cites.
Atlas: Few-shot learning with retrieval augmented language models, 2022
Izacard, G., Lewis, P., Lomeli, M., Hosseini, L., Petroni, F., Schick, T., Dwivedi-Yu, J., Joulin, A., Riedel, S., and Grave, E · 2022
Earlier work this paper cites.
TemporalWiki: A lifelong benchmark for training and evaluating Ever-Evolving language models
Jang, J., Ye, S., Lee, C., Yang, S., Shin, J., Han, J., Kim, G., and Seo, M · 2022
Cited alongside, same era.
StreamingQA: A benchmark for adaptation to new knowledge over time in question answering models
Liška, A., Kočiský, T., Gribovskaya, E., Terzi, T., Sezener, E., Agrawal, D., de Masson d’Autume, C., Scholtes, T., Zaheer, M., Young, S., Gilsenan-McMahon, E., Austin, S., Blunsom, P., and Lazaridou, A · 2022
Cited alongside, same era.
Mass-editing memory in a transformer
Meng, K., Sharma, A. S., Andonian, A., Belinkov, Y., and Bau, D · 2022
Cited alongside, same era.
Memory-based model editing at scale, 2022
Mitchell, E., Lin, C., Bosselut, A., Manning, C. D., and Finn, C · 2022
Cited alongside, same era.
Measuring attribution in natural language generation models, 2022
Rashkin, H., Nikolaev, V., Lamm, M., Aroyo, L., Collins, M., Das, D., Petrov, S., Tomar, G. S., Turc, I., and Reitter, D · 2022
Time sensitive knowledge editing through efficient finetuning
Ge, X., Mousavi, A., Grave, E., Joulin, A., Qian, K., Han, B., Arefiyan, M., and Li, Y · 2024
Later among the works it cites.
Onebench to test them all: Sample-level benchmarking over open-ended capabilities
Ghosh, A., Dziadzio, S., Prabhu, A., Udandarao, V., Albanie, S., and Bethge, M · 2024
Later among the works it cites.
knn-clip: Retrieval enables training-free segmentation on continually expanding large vocabularies
Gui, Z., Sun, S., Li, R., Yuan, J., An, Z., Roth, K., Prabhu, A., and Torr, P · 2024
Later among the works it cites.
Editing the mind of giants: An in-depth exploration of pitfalls of knowledge editing in large language models
Hsueh, C.-H., Huang, P. K.-M., Lin, T.-H., Liao, C.-W., Fang, H.-C., Huang, C.-W., and Chen, Y.-N · 2024
Later among the works it cites.
Simple and scalable strategies to continually pre-train large language models, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Momentum-based weight interpolation of strong zero-shot models for continual learning, 2022
Stojanovski, Z., Roth, K., and Akata, Z · 2022
Cited alongside, same era.
Robust fine-tuning of zero-shot models, 2022
Wortsman, M., Ilharco, G., Kim, J. W., Li, M., Kornblith, S., Roelofs, R., Gontijo-Lopes, R., Hajishirzi, H., Farhadi, A., Namkoong, H., and Schmidt, L · 2022
Cited alongside, same era.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models, 2022
Zaken, E. B., Ravfogel, S., and Goldberg, Y · 2022
Cited alongside, same era.
Disc-medllm: Bridging general large language models and real-world medical consultation
Bao, Z., Chen, W., Xiao, S., Ren, K., Wu, J., Zhong, C., Peng, J., Huang, X., and Wei, Z · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models, 2023
et al., H. T · 2023
Cited alongside, same era.
Aging with GRACE: Lifelong model editing with discrete key-value adaptors
Hartvigsen, T., Sankaranarayanan, S., Palangi, H., Kim, Y., and Ghassemi, M · 2023
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de Las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Cited alongside, same era.
Ibrahim, A., Thérien, B., Gupta, K., Richter, M. L., Anthony, Q., Lesort, T., Belilovsky, E., and Rish, I · 2024
Later among the works it cites.
Medical adaptation of large language and vision-language models: Are we making progress?
Jeong, D. P., Garg, S., Lipton, Z. C., and Oberst, M · 2024
Later among the works it cites.
Khodja, H. A., Béchet, F., Brabant, Q., Nasr, A., and Lecorvé, G · 2024
Later among the works it cites.
Continual learning with weight interpolation, 2024
Kozal, J., Wasilewski, J., Krawczyk, B., and Woźniak, M · 2024
Later among the works it cites.
Large language models in law: A survey
Lai, J., Gan, W., Wu, J., Qi, Z., and Philip, S. Y · 2024
Later among the works it cites.
Language modeling with editable external knowledge, 2024
Li, B. Z., Liu, E., Ross, A., Zeitoun, A., Neubig, G., and Andreas, J · 2024
Later among the works it cites.
Ex-fever: A dataset for multi-hop explainable fact verification, 2024
Ma, H., Xu, W., Wei, Y., Chen, L., Wang, L., Liu, Q., Wu, S., and Wang, L · 2024
Later among the works it cites.
Magmax: Leveraging model merging for seamless continual learning, 2024
Marczak, D., Twardowski, B., Trzciński, T., and Cygert, S · 2024
Later among the works it cites.
Weighted ensemble models are strong continual learners, 2024
Marouf, I. E., Roy, S., Tartaglione, E., and Lathuilière, S · 2024
Later among the works it cites.
Event-level knowledge editing
Peng, H., Wang, X., Li, C., Zeng, K., Duo, J., Cao, Y., Hou, L., and Li, J · 2024
Later among the works it cites.
Lifelong benchmarks: Efficient model evaluation in an era of rapid progress
Prabhu, A., Udandarao, V., Torr, P., Bethge, M., Bibi, A., and Albanie, S · 2024
Later among the works it cites.
Warm: On the benefits of weight averaged reward models, 2024
Ramé, A., Vieillard, N., Hussenot, L., Dadashi, R., Cideron, G., Bachem, O., and Ferret, J · 2024
Later among the works it cites.
A practitioner’s guide to continual multimodal pretraining
Roth, K., Udandarao, V., Dziadzio, S., Prabhu, A., Cherti, M., Vinyals, O., Hénaff, O., Albanie, S., Bethge, M., and Akata, Z · 2024
Later among the works it cites.
Reflecting on the state of rehearsal-free continual learning with pretrained models, 2024
Thede, L., Roth, K., Hénaff, O. J., Bethge, M., and Akata, Z · 2024
Later among the works it cites.
No ”zero-shot” without exponential data: Pretraining concept frequency determines multimodal model performance
Udandarao, V., Prabhu, A., Ghosh, A., Sharma, Y., Torr, P., Bibi, A., Albanie, S., and Bethge, M · 2024
Later among the works it cites.
History matters: Temporal knowledge editing in large language model
Yin, X., Jiang, J., Yang, L., and Wan, X · 2024
Later among the works it cites.
A comprehensive study of knowledge editing for large language models
Zhang, N., Yao, Y., Tian, B., Wang, P., Deng, S., Wang, M., Xi, Z., Mao, S., Zhang, J., Ni, Y., et al · 2024
Later among the works it cites.
Galore: Memory-efficient llm training by gradient low-rank projection, 2024
Zhao, J., Zhang, Z., Chen, B., Wang, Z., Anandkumar, A., and Tian, Y · 2024
Later among the works it cites.
Large language models as reliable knowledge bases?
Zheng, D., Lapata, M., and Pan, J. Z · 2024
Later among the works it cites.
Hipporag: Neurobiologically inspired long-term memory for large language models, 2025
Gutiérrez, B. J., Shu, Y., Gu, Y., Yasunaga, M., and Su, Y · 2025
Closest in time.
Tic-lm: A web-scale benchmark for time-continual llm pretraining, 2025
Li, J., Armandpour, M., Mirzadeh, I., Mehta, S., Shankar, V., Vemulapalli, R., Bengio, S., Tuzel, O., Farajtabar, M., Pouransari, H., and Faghri, F · 2025
Closest in time.