Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have shown great success in various Natural Language Processing (NLP) tasks, whist they still need updates after deployment to fix errors or keep pace with the changing knowledge in the world.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
An estimate of an upper bound for the entropy of English
Brown, P. F.; Della Pietra, S. A.; Della Pietra, V. J.; Lai, J. C.; and Mercer, R. L. 1992 · 1992
Earlier work this paper cites.
Zero-Shot Relation Extraction via Reading Comprehension
Levy, O.; Seo, M.; Choi, E.; and Zettlemoyer, L. 2017 · 2017
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; De Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; and Gelly, S. 2019 · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Kwiatkowski, T.; Palomaki, J.; Redfield, O.; Collins, M.; Parikh, A.; Alberti, C.; Epstein, D.; Polosukhin, I.; Devlin, J.; Lee, K.; et al. 2019 · 2019
Earlier work this paper cites.
Transformer Feed-Forward Layers Are Key-Value Memories
Geva, M.; Schuster, R.; Berant, J.; and Levy, O. 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Earlier work this paper cites.
Prefix-Tuning: Optimizing Continuous Prompts for Generation
Li, X. L.; and Liang, P. 2021 · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R.; Hilton, J.; Balaji, S.; Wu, J.; Ouyang, L.; Kim, C.; Hesse, C.; Jain, S.; Kosaraju, V.; Saunders, W.; et al. 2021 · 2021
Earlier work this paper cites.
Spot: Better frozen model adaptation through soft prompt transfer
Vu, T.; Lester, B.; Constant, N.; Al-Rfou, R.; and Cer, D. 2021 · 2021
Cited alongside, same era.
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models
Ben Zaken, E.; Goldberg, Y.; and Ravfogel, S. 2022 · 2022
Cited alongside, same era.
FairLex: A multilingual benchmark for evaluating fairness in legal text processing
Chalkidis, I.; Pasini, T.; Zhang, S.; Tomada, L.; Schwemer, S. F.; and Søgaard, A. 2022 · 2022
Cited alongside, same era.
Calibrating Factual Knowledge in Pretrained Language Models
Dong, Q.; Dai, D.; Song, Y.; Xu, J.; Sui, Z.; and Li, L. 2022 · 2022
Cited alongside, same era.
Aging with GRACE: Lifelong Model Editing with Discrete Key-Value Adaptors
Hartvigsen, T.; Sankaranarayanan, S.; Palangi, H.; Kim, Y.; and Ghassemi, M. 2022 · 2022
Transformer-Patcher: One Mistake worth One Neuron
Huang, Z.; Shen, Y.; Zhang, X.; Zhou, J.; Rong, W.; and Xiong, Z. 2023 · 2023
Closest in time.
Beyond One-Model-Fits-All: A Survey of Domain Specialization for Large Language Models
Ling, C.; Zhao, X.; Lu, J.; Deng, C.; Zheng, C.; Wang, J.; Chowdhury, T.; Li, Y.; Cui, H.; Zhao, T.; et al. 2023 · 2023
Closest in time.
Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models
Manakul, P.; Liusie, A.; and Gales, M. J. 2023 · 2023
Closest in time.
Toolformer: Language models can teach themselves to use tools
Schick, T.; Dwivedi-Yu, J.; Dessì, R.; Raileanu, R.; Lomeli, M.; Zettlemoyer, L.; Cancedda, N.; and Scialom, T. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Visual Prompt Tuning
Jia, M.; Tang, L.; Chen, B.-C.; Cardie, C.; Belongie, S.; Hariharan, B.; and Lim, S.-N. 2022 · 2022
Cited alongside, same era.
On Continual Model Refinement in Out-of-Distribution Data Streams
Lin, B. Y.; Wang, S.; Lin, X. V.; Jia, R.; Xiao, L.; Ren, X.; and Yih, W.-t. 2022 · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q. V.; Zhou, D.; et al. 2022 · 2022
Cited alongside, same era.
Locating and editing factual associations in GPT
Meng, K.; Bau, D.; Andonian, A.; and Belinkov, Y. 2022a
Cited in the paper.
Mass-editing memory in a transformer
Meng, K.; Sharma, A. S.; Andonian, A.; Belinkov, Y.; and Bau, D. 2022b
Cited in the paper.
Fast Model Editing at Scale
Mitchell, E.; Lin, C.; Bosselut, A.; Finn, C.; and Manning, C. D. 2022a
Cited in the paper.
Memory-based model editing at scale
Mitchell, E.; Lin, C.; Bosselut, A.; Manning, C. D.; and Finn, C. 2022b
Cited in the paper.
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Closest in time.
DyLoRA: Parameter-Efficient Tuning of Pre-trained Models using Dynamic Search-Free Low-Rank Adaptation
Valipour, M.; Rezagholizadeh, M.; Kobyzev, I.; and Ghodsi, A. 2023 · 2023
Closest in time.
Editing Large Language Models: Problems, Methods, and Opportunities
Yao, Y.; Wang, P.; Tian, B.; Cheng, S.; Li, Z.; Deng, S.; Chen, H.; and Zhang, N. 2023 · 2023
Closest in time.
Adaptive budget allocation for parameter-efficient fine-tuning
Zhang, Q.; Chen, M.; Bukharin, A.; He, P.; Cheng, Y.; Chen, W.; and Zhao, T. 2023 · 2023
Closest in time.