Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are not amenable to frequent re-training, due to high training costs arising from their massive scale.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, et al · 2017
Earlier work this paper cites.
Gradient episodic memory for continual learning
David Lopez-Paz and Marc’Aurelio Ranzato · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Continual learning with deep generative replay
Hanul Shin, Jung Kwon Lee, Jaehong Kim, and Jiwon Kim · 2017
Earlier work this paper cites.
Piggyback: Adapting a single network to multiple tasks by learning to mask weights
Arun Mallya, Dillon Davis, and Svetlana Lazebnik · 2018
Earlier work this paper cites.
On tiny episodic memories in continual learning
Arslan Chaudhry, Marcus Rohrbach, Mohamed Elhoseiny, et al · 2019
Earlier work this paper cites.
Unified language model pre-training for natural language understanding and generation
Li Dong, Nan Yang, Wenhui Wang, et al · 2019
Earlier work this paper cites.
Continual lifelong learning in natural language processing: A survey
Magdalena Biesialska, Katarzyna Biesialska, and Marta R. Costa-jussà · 2020
Earlier work this paper cites.
Artificial Intelligence, Values, and Alignment
Iason Gabriel · 2020
Earlier work this paper cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, et al · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive NLP tasks
Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, et al · 2020
Earlier work this paper cites.
ERNIE 2.0: A continual pre-training framework for language understanding
Yu Sun, Shuohuan Wang, Yu-Kun Li, et al · 2020
Earlier work this paper cites.
Learning to solve NLP tasks in an incremental number of languages
Giuseppe Castellucci, Simone Filice, Danilo Croce, and Roberto Basili · 2021
Earlier work this paper cites.
Curriculum-meta learning for order-robust continual relation extraction
Tongtong Wu, Xuekai Li, Yuan-Fang Li, Gholamreza Haffari, Guilin Qi, Yujin Zhu, and Guoqiang Xu · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, et al · 2022
Earlier work this paper cites.
Continual pre-training mitigates forgetting in language and vision
Andrea Cossu, Tinne Tuytelaars, Antonio Carta, et al · 2022
Earlier work this paper cites.
Understanding dataset difficulty with V -usable information
Kawin Ethayarajh, Yejin Choi, and Swabha Swayamdipta · 2022
Earlier work this paper cites.
Temporalwiki: A lifelong benchmark for training and evaluating ever-evolving language models
Joel Jang, Seonghyeon Ye, Changho Lee, et al · 2022
Earlier work this paper cites.
Towards continual knowledge learning of language models
Joel Jang, Seonghyeon Ye, Sohee Yang, et al · 2022
Earlier work this paper cites.
Lifelong pretraining: Continually adapting language models to emerging corpora
Xisen Jin, Dejiao Zhang, Henghui Zhu, et al · 2022
Earlier work this paper cites.
Fine-tuned language models are continual learners
Thomas Scialom, Tuhin Chakrabarty, and Smaranda Muresan · 2022
Earlier work this paper cites.
Pretrained language model in continual learning: A comparative study
Tongtong Wu, Massimo Caccia, Zhuang Li, Yuan-Fang Li, Guilin Qi, and Gholamreza Haffari · 2022
Earlier work this paper cites.
Contintin: Continual learning from task instructions
Wenpeng Yin, Jia Li, and Caiming Xiong · 2022
Earlier work this paper cites.
CERT: continual pre-training on sketches for library-oriented code generation
Daoguang Zan, Bei Chen, Dejian Yang, et al · 2022
Earlier work this paper cites.
Llemma: An open language model for mathematics
Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, et al · 2023
Cited alongside, same era.
Unlearn what you want to forget: Efficient unlearning for llms
Jiaao Chen and Diyi Yang · 2023
Cited alongside, same era.
Continual multimodal knowledge graph construction
Xiang Chen, Jintian Zhang, Xiaohan Wang, et al · 2023
Cited alongside, same era.
Learn from yesterday: A semi-supervised continual learning method for supervision-limited text-to-sql task streams
Yongrui Chen, Xinnan Guo, Tongtong Wu, et al · 2023
Cited alongside, same era.
Adapting large language models via reading comprehension
Daixuan Cheng, Shaohan Huang, and Furu Wei · 2023
Cited alongside, same era.
Language model with plug-in knowledge memory, 2023
Progressive prompts: Continual learning for language models
Anastasia Razdaibiedina, Yuning Mao, Rui Hou, et al · 2023
Later among the works it cites.
Conpet: Continual parameter-efficient tuning for large language models
Chenyang Song, Xu Han, Zheni Zeng, et al · 2023
Later among the works it cites.
Continual learning for instruction following from realtime feedback
Alane Suhr and Yoav Artzi · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, et al · 2023
Later among the works it cites.
Continual learning: Applications and the road forward
Eli Verwimp, Rahaf Aljundi, Shai Ben-David, Matthias Bethge, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xin Cheng, Yankai Lin, Dongyan Zhao, and Rui Yan · 2023
Cited alongside, same era.
How abilities in large language models are affected by supervised fine-tuning data composition
Guanting Dong, Hongyi Yuan, Keming Lu, et al · 2023
Cited alongside, same era.
A study of continual learning under language shift
Evangelia Gogoulou, Timothée Lesort, Magnus Boman, and Joakim Nivre · 2023
Cited alongside, same era.
Toolkengpt: Augmenting frozen language models with massive tools via tool embeddings
Shibo Hao, Tianyang Liu, Zhen Wang, and Zhiting Hu · 2023
Cited alongside, same era.
Drinking from a firehose: Continual learning with web-scale natural language
Hexiang Hu, Ozan Sener, Fei Sha, and Vladlen Koltun · 2023
Cited alongside, same era.
Exploring the benefits of training expert language models over instruction tuning
Joel Jang, Seungone Kim, Seonghyeon Ye, et al · 2023
Cited alongside, same era.
Qiao Jin, Yifan Yang, Qingyu Chen, and Zhiyong Lu · 2023
Cited alongside, same era.
Orthogonal subspace learning for language model continual learning
Xiao Wang, Tianze Chen, Qiming Ge, et al · 2023
Later among the works it cites.
Trace: A comprehensive benchmark for continual learning in large language models
Xiao Wang, Yuansen Zhang, Tianze Chen, et al · 2023
Later among the works it cites.
Aligning large language models with human: A survey
Yufei Wang, Wanjun Zhong, Liangyou Li, et al · 2023
Later among the works it cites.
Efficient continual pre-training for building domain specific large language models
Yong Xie, Karan Aggarwal, and Aitzaz Ahmad · 2023
Later among the works it cites.
Rationale-enhanced language models are better continual relation learners
Weimin Xiong, Yifan Song, Peiyi Wang, and Sujian Li · 2023
Later among the works it cites.
Exploring continual learning for code generation models
Prateek Yadav, Qing Sun, Hantian Ding, et al · 2023
Later among the works it cites.
Editing large language models: Problems, methods, and opportunities
Yunzhi Yao, Peng Wang, Bozhong Tian, et al · 2023
Later among the works it cites.
Copf: Continual learning human preference through optimal policy fitting
Han Zhang, Lin Gui, Yuanzhao Zhai, et al · 2023
Later among the works it cites.
Instruction tuning for large language models: A survey
Shengyu Zhang, Linfeng Dong, Xiaoya Li, et al · 2023
Later among the works it cites.
Reformulating domain adaptation of large language models as adapt-retrieve-revise
Yating Zhang, Yexiang Wang, Fei Cheng, et al · 2023
Later among the works it cites.
How do large language models capture the ever-changing world knowledge? A review of recent advances
Zihan Zhang, Meng Fang, Ling Chen, et al · 2023
Later among the works it cites.
Citb: A benchmark for continual instruction tuning
Zihan Zhang, Meng Fang, Ling Chen, and Mohammad-Reza Namazi-Rad · 2023
Later among the works it cites.
CPPO: Continual learning for reinforcement learning with human feedback
Anonymous · 2024
Closest in time.
Scalable language model with generalized continual learning
Anonymous · 2024
Closest in time.
Towards lifelong scene graph generation with knowledge-ware in-context prompt learning
Tao He, Tongtong Wu, Dongyang Zhang, Guiduo Duan, Ke Qin, and Yuan-Fang Li · 2024
Closest in time.
Autoact: Automatic agent learning from scratch via self-planning
Shuofei Qiao, Ningyu Zhang, Runnan Fang, Yujie Luo, Wangchunshu Zhou, Yuchen Eleanor Jiang, Chengfei Lv, and Huajun Chen · 2024
Closest in time.
Llama pro: Progressive llama with block expansion
Chengyue Wu, Yukang Gan, Yixiao Ge, et al · 2024
Closest in time.
Pllama: An open-source large language model for plant science
Xianjun Yang, Junfeng Gao, Wenxin Xue, and Erik Alexandersson · 2024
Closest in time.
Dapt: A dual attention framework for parameter-efficient continual learning of large language models
Weixiang Zhao, Shilong Wang, Yulin Hu, et al · 2024
Closest in time.