Fetching the paper…
Reading the bibliography…
As the applications of large language models (LLMs) expand across diverse fields, the ability of these models to adapt to ongoing changes in data, tasks, and user preferences becomes crucial.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Decomposing logits distillation for incremental named entity recognition. In Proceedings of International ACM SIGIR Conference on Research and Development in Information Retrieval . 1919–1923
Duzhen Zhang, Yahan Yu, Feilong Chen, and Xiuyi Chen. 2023f · 1923
Earlier work this paper cites.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning
Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin A Raffel. 2022b · 1965
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen. 1989 · 1989
Earlier work this paper cites.
OntoNotes: the 90% solution. In Proceedings of the human language technology conference of the NAACL . 57–60
Eduard Hovy, Mitch Marcus, Martha Palmer, Lance Ramshaw, and Ralph Weischedel. 2006 · 2006
Earlier work this paper cites.
Causal inference in statistics: An overview
Judea Pearl. 2009 · 2009
Earlier work this paper cites.
Serving the enterprise and beyond with informatics for integrating biology and the bedside (i2b2)
Shawn N Murphy, Griffin Weber, Michael Mendis, Vivian Gainer, Henry C Chueh, Susanne Churchill, and Isaac Kohane. 2010 · 2010
Earlier work this paper cites.
The stability-plasticity dilemma: Investigating the continuum from catastrophic forgetting to age-limited learning effects
Martial Mermillod, Aurélia Bugaiska, and Patrick Bonin. 2013 · 2013
Earlier work this paper cites.
Net2net: Accelerating learning via knowledge transfer
Tianqi Chen, Ian Goodfellow, and Jonathon Shlens. 2015 · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. In International Conference on Learning Representations
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2016 · 2016
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Earlier work this paper cites.
Learning without forgetting
Zhizhong Li and Derek Hoiem. 2017 · 2017
Earlier work this paper cites.
Continual learning with deep generative replay
Hanul Shin, Jung Kwon Lee, Jaehong Kim, and Jiwon Kim. 2017 · 2017
Earlier work this paper cites.
Tl; dr: Mining reddit to learn automatic summarization. In Proceedings of the Workshop on New Frontiers in Summarization . 59–63
Michael Völske, Martin Potthast, Shahbaz Syed, and Benno Stein. 2017 · 2017
Earlier work this paper cites.
Position-aware Attention and Supervised Data Improve Slot Filling. In Proceedings of Empirical Methods in Natural Language Processing . 35–45
Yuhao Zhang, Victor Zhong, Danqi Chen, Gabor Angeli, and Christopher D Manning. 2017 · 2017
Earlier work this paper cites.
Memory aware synapses: Learning what (not) to forget. In Proceedings of the European conference on computer vision . 139–154
Rahaf Aljundi, Francesca Babiloni, Mohamed Elhoseiny, Marcus Rohrbach, and Tinne Tuytelaars. 2018 · 2018
Earlier work this paper cites.
Lifelong machine learning . Vol. 1
Zhiyuan Chen and Bing Liu. 2018 · 2018
Earlier work this paper cites.
FewRel: A Large-Scale Supervised Few-Shot Relation Classification Dataset with State-of-the-Art Evaluation. In Proceedings of Empirical Methods in Natural Language Processing . 4803–4809
Xu Han, Hao Zhu, Pengfei Yu, Ziyun Wang, Yuan Yao, Zhiyuan Liu, and Maosong Sun. 2018 · 2018
Earlier work this paper cites.
Regularized training objective for continued training for domain adaptation in neural machine translation. In Proceedings of Workshop on Neural Machine Translation and Generation . 36–44
Huda Khayrallah, Brian Thompson, Kevin Duh, and Philipp Koehn. 2018 · 2018
Earlier work this paper cites.
On first-order meta-learning algorithms
Alex Nichol, Joshua Achiam, and John Schulman. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
FEVER: a Large-scale Dataset for Fact Extraction and VERification. In Proceedings of Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . 809–819
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018 · 2018
Earlier work this paper cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. In International Conference on Learning Representations
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018 · 2018
Earlier work this paper cites.
From Bilingual to Multilingual Neural Machine Translation by Incremental Training. In Proceedings of Annual Meeting of the Association for Computational Linguistics: Student Research Workshop . 236–242
Carlos Escolano, Marta R Costa-jussà, and José AR Fonollosa. 2019 · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP. In International conference on machine learning . 2790–2799
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Earlier work this paper cites.
An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction. In Proceedings of Empirical Methods in Natural Language Processing and the International Joint Conference on Natural Language Processing . 1311–1316
Stefan Larson, Anish Mahendran, Joseph J Peper, Christopher Clarke, Andrew Lee, Parker Hill, Jonathan K Kummerfeld, Kevin Leach, Michael A Laurenzano, Lingjia Tang, et al · 2019
Earlier work this paper cites.
Meta-Learning Improves Lifelong Relation Extraction. In Proceedings of the 4th Workshop on Representation Learning for NLP . 224–229
Abiola Obamuyide and Andreas Vlachos. 2019 · 2019
Earlier work this paper cites.
Continual lifelong learning with neural networks: A review
German I Parisi, Ronald Kemker, Jose L Part, Christopher Kanan, and Stefan Wermter. 2019 · 2019
Earlier work this paper cites.
A progressive model to enable continual learning for semantic slot filling. In Proceedings of the Conference on Empirical Methods in Natural Language Processing and the International Joint Conference on Natural Language Processing . 1279–1284
Yilin Shen, Xiangyu Zeng, and Hongxia Jin. 2019 · 2019
Earlier work this paper cites.
LAMOL: LAnguage MOdeling for Lifelong Language Learning. In International Conference on Learning Representations
Fan-Keng Sun, Cheng-Hao Ho, and Hung-Yi Lee. 2019 · 2019
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019a · 2019
Earlier work this paper cites.
Scalable and Order-robust Continual Learning with Additive Parameter Decomposition. In International Conference on Learning Representations
Jaehong Yoon, Saehoon Kim, Eunho Yang, and Sung Ju Hwang. 2019 · 2019
Earlier work this paper cites.
Rapid trial-and-error learning with simulation supports flexible tool use and physical reasoning
Kelsey R Allen, Kevin A Smith, and Joshua B Tenenbaum. 2020 · 2020
Earlier work this paper cites.
Findings of the first shared task on lifelong learning machine translation. In EMNLP 2020, Fifth Conference on Machine Translation . 56–64
Loïc Barrault, Magdalena Marta Biesialska, Marta Ruiz Costa-Jussà, Fethi Bougares, and Olivier Galibert. 2020 · 2020
Earlier work this paper cites.
Continual Lifelong Learning in Natural Language Processing: A Survey. In Proceedings of International Conference on Computational Linguistics . 6523–6541
Magdalena Biesialska, Katarzyna Biesialska, and Marta R Costa-jussà. 2020 · 2020
Earlier work this paper cites.
Incremental event detection via knowledge consolidation networks. In Proceedings of Conference on Empirical Methods in Natural Language Processing . 707–717
Pengfei Cao, Yubo Chen, Jun Zhao, and Taifeng Wang. 2020 · 2020
Earlier work this paper cites.
Efficient Intent Detection with Dual Sentence Encoders. In Proceedings of the 2nd Workshop on Natural Language Processing for Conversational AI . 38–45
Iñigo Casanueva, Tadas Temčinas, Daniela Gerz, Matthew Henderson, and Ivan Vulić. 2020 · 2020
Earlier work this paper cites.
Recall and Learn: Fine-tuning Deep Pretrained Language Models with Less Forgetting. In Proceedings of Conference on Empirical Methods in Natural Language Processing . 7870–7881
Sanyuan Chen, Yutai Hou, Yiming Cui, Wanxiang Che, Ting Liu, and Xiangzhan Yu. 2020 · 2020
Earlier work this paper cites.
Don’t Stop Pretraining: Adapt Language Models to Domains and Tasks. In Proceedings of Annual Meeting of the Association for Computational Linguistics . 8342–8360
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith. 2020 · 2020
Earlier work this paper cites.
Econet: Effective continual pretraining of language models for event temporal reasoning
Rujun Han, Xiang Ren, and Nanyun Peng. 2020b · 2020
Earlier work this paper cites.
A survey on contrastive self-supervised learning
Ashish Jaiswal, Ashwin Ramesh Babu, Mohammad Zaki Zadeh, Debapriya Banerjee, and Fillia Makedon. 2020 · 2020
Earlier work this paper cites.
Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of Conference on Empirical Methods in Natural Language Processing . 6769–6781
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020 · 2020
Earlier work this paper cites.
Continual learning for robotics: Definition, framework, learning strategies, opportunities and challenges
Timothée Lesort, Vincenzo Lomonaco, Andrei Stoian, Davide Maltoni, David Filliat, and Natalia Díaz-Rodríguez. 2020 · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
Continual Learning for Natural Language Generation in Task-oriented Dialog Systems. In Findings of EMNLP . 3461–3474
Fei Mi, Liangwei Chen, Mengjie Zhao, Minlie Huang, and Boi Faltings. 2020 · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020 · 2020
Earlier work this paper cites.
Distill and replay for continual language learning. In Proceedings of international conference on computational linguistics . 3569–3579
Jingyuan Sun, Shaonan Wang, Jiajun Zhang, and Chengqing Zong. 2020 · 2020
Earlier work this paper cites.
CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data. In Proceedings of the Twelfth Language Resources and Evaluation Conference . 4003–4012
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Édouard Grave. 2020 · 2020
Earlier work this paper cites.
Continual Learning in Multilingual NMT via Language-Specific Embeddings. In Proceedings of the Sixth Conference on Machine Translation . 542–565
Alexandre Bérard. 2021 · 2021
Earlier work this paper cites.
Continual learning for neural machine translation. In Proceedings of Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . 3964–3974
Yue Cao, Hao-Ran Wei, Boxing Chen, and Xiaojun Wan. 2021 · 2021
Earlier work this paper cites.
Learning to solve NLP tasks in an incremental number of languages. In Proceedings of Annual Meeting of the Association for Computational Linguistics and International Joint Conference on Natural Language Processing . 837–847
Giuseppe Castellucci, Simone Filice, Danilo Croce, and Roberto Basili. 2021 · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Refining sample embeddings with relation prototypes to enhance continual relation extraction. In Proceedings of Annual Meeting of the Association for Computational Linguistics and International Joint Conference on Natural Language Processing . 232–243
Li Cui, Deqing Yang, Jiaxin Yu, Chengwei Hu, Jiayang Cheng, Jingjie Yi, and Yanghua Xiao. 2021 · 2021
Earlier work this paper cites.
Knowledge neurons in pretrained transformers
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2021 · 2021
Earlier work this paper cites.
Editing factual knowledge in language models
Nicola De Cao, Wilker Aziz, and Ivan Titov. 2021 · 2021
Earlier work this paper cites.
A continual learning survey: Defying forgetting in classification tasks
Matthias De Lange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia, Aleš Leonardis, Gregory Slabaugh, and Tinne Tuytelaars. 2021 · 2021
Earlier work this paper cites.
Few-NERD: A Few-shot Named Entity Recognition Dataset. In Proceedings of Annual Meeting of the Association for Computational Linguistics and International Joint Conference on Natural Language Processing . 3198–3213
Ning Ding, Guangwei Xu, Yulin Chen, Xiaobin Wang, Xu Han, Pengjun Xie, Haitao Zheng, and Zhiyuan Liu. 2021 · 2021
Earlier work this paper cites.
Towards Continual Learning for Multilingual Machine Translation via Vocabulary Substitution. In Proceedings of Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . 1184–1192
Xavier Garcia, Noah Constant, Ankur Parikh, and Orhan Firat. 2021 · 2021
Earlier work this paper cites.
Analyzing the forgetting problem in pretrain-finetuning of open-domain dialogue response models. In Proceedings of Conference of the European Chapter of the Association for Computational Linguistics . 1121–1133
Tianxing He, Jun Liu, Kyunghyun Cho, Myle Ott, Bing Liu, James Glass, and Fuchun Peng. 2021 · 2021
Earlier work this paper cites.
LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations
Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al · 2021
Earlier work this paper cites.
Hyperparameter-free Continuous Learning for Domain Classification in Natural Language Understanding. In Proceedings of Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . 2669–2678
Ting Hua, Yilin Shen, Changsheng Zhao, Yen-Chang Hsu, and Hongxia Jin. 2021 · 2021
Earlier work this paper cites.
Continual Learning for Text Classification with Information Disentanglement Based Regularization. In Proceedings of Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . 2736–2746
Yufan Huang, Yanzhe Zhang, Jiaao Chen, Xuezhi Wang, and Diyi Yang. 2021 · 2021
Earlier work this paper cites.
Learn Continually, Generalize Rapidly: Lifelong Knowledge Accumulation for Few-shot Learning. In Findings of EMNLP . 714–729
Xisen Jin, Bill Yuchen Lin, Mohammad Rostami, and Xiang Ren. 2021 · 2021
Earlier work this paper cites.
Rational LAMOL: A rationale-based lifelong learning framework. In Proceedings of Annual Meeting of the Association for Computational Linguistics and International Joint Conference on Natural Language Processing . 2942–2953
Kasidis Kanwatchara, Thanapapas Horsuwan, Piyawat Lertvittayakumjorn, Boonserm Kijsirikul, and Peerapon Vateekul. 2021 · 2021
Earlier work this paper cites.
Achieving forgetting prevention and knowledge transfer in continual learning
Zixuan Ke, Bing Liu, Nianzu Ma, Hu Xu, and Lei Shu. 2021a · 2021
Earlier work this paper cites.
Mind the gap: Assessing temporal generalization in neural language models
Angeliki Lazaridou, Adhi Kuncoro, Elena Gribovskaya, Devang Agrawal, Adam Liska, Tayfun Terzi, Mai Gimenez, Cyprien de Masson d’Autume, Tomas Kocisky, Sebastian Ruder, et al · 2021
Earlier work this paper cites.
Sequential Reptile: Inter-Task Gradient Alignment for Multilingual Learning. In International Conference on Learning Representations
Seanie Lee, Hae Beom Lee, Juho Lee, and Sung Ju Hwang. 2021 · 2021
Earlier work this paper cites.
The Power of Scale for Parameter-Efficient Prompt Tuning. In Proceedings of Empirical Methods in Natural Language Processing . 3045–3059
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Earlier work this paper cites.
Prefix-Tuning: Optimizing Continuous Prompts for Generation. In Proceedings of Annual Meeting of the Association for Computational Linguistics and International Joint Conference on Natural Language Processing . 4582–4597
Xiang Lisa Li and Percy Liang. 2021 · 2021
Earlier work this paper cites.
Lifelong intent detection via multi-strategy rebalancing
Qingbin Liu, Xiaoyan Yu, Shizhu He, Kang Liu, and Jun Zhao. 2021c · 2021
Earlier work this paper cites.
Continual Learning in Task-Oriented Dialogue Systems. In Proceedings of Conference on Empirical Methods in Natural Language Processing . 7452–7467
Andrea Madotto, Zhaojiang Lin, Zhenpeng Zhou, Seungwhan Moon, Paul A Crook, Bing Liu, Zhou Yu, Eunjoon Cho, Pascale Fung, and Zhiguang Wang. 2021 · 2021
Earlier work this paper cites.
Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D Manning. 2021 · 2021
Earlier work this paper cites.
Continual learning for named entity recognition. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 35. 13570–13577
Natawut Monaikul, Giuseppe Castellucci, Simone Filice, and Oleg Rokhlenko. 2021 · 2021
Earlier work this paper cites.
Continual few-shot learning for text classification. In Proceedings of Conference on Empirical Methods in Natural Language Processing . 5688–5702
Ramakanth Pasunuru, Veselin Stoyanov, and Mohit Bansal. 2021 · 2021
Earlier work this paper cites.
Lifelong Learning of Hate Speech Classification on Social Media. In Proceedings of Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . 2304–2314
Jing Qian, Hong Wang, Mai ElSherief, and Xifeng Yan. 2021 · 2021
Earlier work this paper cites.
LFPT5: A Unified Framework for Lifelong Few-shot Language Learning Based on Prompt Tuning of T5. In International Conference on Learning Representations
Chengwei Qin and Shafiq Joty. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision. In International conference on machine learning . 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Temporal Adaptation of BERT and Performance on Downstream Document Classification: Insights from Social Media. In Findings of EMNLP 2021 . 2400–2412
Paul Röttger and Janet Pierrehumbert. 2021 · 2021
Cited alongside, same era.
Lifelong knowledge-enriched social event representation learning. In Proceedings of Conference of the European Chapter of the Association for Computational Linguistics . 3624–3635
Prashanth Vijayaraghavan and Deb Roy. 2021 · 2021
Cited alongside, same era.
Moduleformer: Learning modular large language models from uncurated data
Yikang Shen, Zheyu Zhang, Tianyou Cao, Shawn Tan, Zhenfang Chen, and Chuang Gan. 2023b · 2023
Later among the works it cites.
Large language models encode clinical knowledge
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al · 2023
Later among the works it cites.
Continual diffusion: Continual customization of text-to-image diffusion with c-lora
James Seale Smith, Yen-Chang Hsu, Lingyu Zhang, Ting Hua, Zsolt Kira, Yilin Shen, and Hongxia Jin. 2023a · 2023
Later among the works it cites.
Conpet: Continual parameter-efficient tuning for large language models
Chenyang Song, Xu Han, Zheni Zeng, Kuai Li, Chen Chen, Zhiyuan Liu, Maosong Sun, and Tao Yang. 2023a · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Curriculum-meta learning for order-robust continual relation extraction. In Proceedings of the AAAI conference on artificial intelligence , Vol. 35. 10363–10369
Tongtong Wu, Xuekai Li, Yuan-Fang Li, Gholamreza Haffari, Guilin Qi, Yujin Zhu, and Guoqiang Xu. 2021 · 2021
Cited alongside, same era.
Incremental Few-shot Text Classification with Multi-round New Classes: Formulation, Dataset and System. In Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . 1351–1360
Congying Xia, Wenpeng Yin, Yihao Feng, and Philip Yu. 2021 · 2021
Cited alongside, same era.
Lifelong event detection with knowledge transfer. In Proceedings of Conference on Empirical Methods in Natural Language Processing . 5278–5290
Pengfei Yu, Heng Ji, and Prem Natarajan. 2021 · 2021
Cited alongside, same era.
Incremental intent detection for medical domain with contrast replay networks. In Findings of ACL 2022 . 3549–3556
Guirong Bai, Shizhu He, Kang Liu, and Jun Zhao. 2022 · 2022
Cited alongside, same era.
Continual pre-training mitigates forgetting in language and vision
Andrea Cossu, Tinne Tuytelaars, Antonio Carta, Lucia Passaro, Vincenzo Lomonaco, and Davide Bacciu. 2022 · 2022
Cited alongside, same era.
Memory efficient continual learning with transformers
Beyza Ermis, Giovanni Zappella, Martin Wistuba, Aditya Rawal, and Cedric Archambeau. 2022 · 2022
Cited alongside, same era.
Continual Learning of Neural Machine Translation within Low Forgetting Risk Regions. In Proceedings of Conference on Empirical Methods in Natural Language Processing . 1707–1718
Shuhao Gu, Bojie Hu, and Yang Feng. 2022 · 2022
Cited alongside, same era.
DEMix Layers: Disentangling Domains for Modular Language Modeling. In Proceedings of Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . 5557–5576
Suchin Gururangan, Mike Lewis, Ari Holtzman, Noah A Smith, and Luke Zettlemoyer. 2022 · 2022
Cited alongside, same era.
InfoCL: Alleviating Catastrophic Forgetting in Continual Text Classification from An Information Theoretic Perspective. In Findings of EMNLP 2023 . 14557–14570
Yifan Song, Peiyi Wang, Weimin Xiong, Dawei Zhu, Tianyu Liu, Zhifang Sui, and Sujian Li. 2023b · 2023
Later among the works it cites.
A Simple and Effective Pruning Approach for Large Language Models. In The Twelfth International Conference on Learning Representations
Mingjie Sun, Zhuang Liu, Anna Bair, and J Zico Kolter. 2023 · 2023
Later among the works it cites.
Toolalpaca: Generalized tool learning for language models with 3000 simulated cases
Qiaoyu Tang, Ziliang Deng, Hongyu Lin, Xianpei Han, Qiao Liang, and Le Sun. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions. In Proceedings of Annual Meeting of the Association for Computational Linguistics . 10014–10037
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2023 · 2023
Later among the works it cites.
Knowledge editing for large language models: A survey
Song Wang, Yaochen Zhu, Haochen Liu, Zaiyi Zheng, Chen Chen, et al · 2023
Later among the works it cites.
Orthogonal Subspace Learning for Language Model Continual Learning. In Findings of EMNLP 2023 . 10658–10671
Xiao Wang, Tianze Chen, Qiming Ge, Han Xia, Rong Bao, Rui Zheng, Qi Zhang, Tao Gui, and Xuan-Jing Huang. 2023a · 2023
Later among the works it cites.
Serial Contrastive Knowledge Distillation for Continual Few-shot Relation Extraction. In Findings of ACL 2023 . 12693–12706
Xinyi Wang, Zitao Wang, and Wei Hu. 2023d · 2023
Later among the works it cites.
Overcoming Catastrophic Forgetting in Massively Multilingual Continual Learning. In Findings of ACL 2023 . 768–777
Genta Winata, Lingjue Xie, Karthik Radhakrishnan, Shijie Wu, Xisen Jin, Pengxiang Cheng, Mayank Kulkarni, and Daniel Preoţiuc-Pietro. 2023 · 2023
Later among the works it cites.
Continual Learning with Low Rank Adaptation. In NeurIPS 2023 Workshop on Distribution Shifts: New Frontiers with Foundation Models
Martin Wistuba, Lukas Balles, Giovanni Zappella, et al · 2023
Later among the works it cites.
Enhancing Continual Relation Extraction via Classifier Decomposition. In Findings of ACL 2023 . 10053–10062
Heming Xia, Peiyi Wang, Tianyu Liu, Binghuai Lin, Yunbo Cao, and Zhifang Sui. 2023 · 2023
Later among the works it cites.
Efficient continual pre-training for building domain specific large language models
Yong Xie, Karan Aggarwal, and Aitzaz Ahmad. 2023a · 2023
Later among the works it cites.
Exploring Continual Learning for Code Generation Models. In Proceedings of Annual Meeting of the Association for Computational Linguistics . 782–792
Prateek Yadav, Qing Sun, Hantian Ding, Xiaopeng Li, Dejiao Zhang, Ming Tan, Parminder Bhatia, Xiaofei Ma, Ramesh Nallapati, Murali Krishna Ramanathan, et al · 2023
Later among the works it cites.
FinGPT: Open-Source Financial Large Language Models
Hongyang Yang, Xiao-Yang Liu, and Christina Dan Wang. 2023 · 2023
Later among the works it cites.
A Survey on Continual Semantic Segmentation: Theory, Challenge, Method and Application
Bo Yuan and Danpei Zhao. 2023 · 2023
Later among the works it cites.
Continual graph learning: A survey
Qiao Yuan, Sheng-Uei Guan, Pin Ni, Tianlun Luo, Ka Lok Man, Prudence Wong, and Victor Chang. 2023 · 2023
Later among the works it cites.
Removing rlhf protections in gpt-4 via fine-tuning
Qiusi Zhan, Richard Fang, Rohan Bindu, Akul Gupta, Tatsunori Hashimoto, and Daniel Kang. 2023 · 2023
Later among the works it cites.
Mitigating Temporal Misalignment by Discarding Outdated Facts. In Proceedings of Conference on Empirical Methods in Natural Language Processing . 14213–14226
Michael Zhang and Eunsol Choi. 2023 · 2023
Later among the works it cites.
A neural span-based continual named entity recognition model. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 37. 13993–14001
Yunan Zhang and Qingcai Chen. 2023 · 2023
Later among the works it cites.
Learning and forgetting unsafe examples in large language models
Jiachen Zhao, Zhun Deng, David Madras, James Zou, and Mengye Ren. 2023b · 2023
Later among the works it cites.
Inrank: Incremental low-rank learning
Jiawei Zhao, Yifei Zhang, Beidi Chen, Florian Schäfer, and Anima Anandkumar. 2023c · 2023
Later among the works it cites.
Learn or Recall? Revisiting Incremental Learning with Pre-trained Language Models
Junhao Zheng, Shengjie Qiu, and Qianli Ma. 2023b · 2023
Later among the works it cites.
A Survey on Data Selection for Language Models
Alon Albalak, Yanai Elazar, Sang Michael Xie, Shayne Longpre, Nathan Lambert, Xinyi Wang, Niklas Muennighoff, Bairu Hou, Liangming Pan, Haewon Jeong, et al · 2024
Closest in time.
CoIN: A Benchmark of Continual Instruction tuNing for Multimodel Large Language Model
Cheng Chen, Junchen Zhu, Xu Luo, Hengtao Shen, Lianli Gao, and Jingkuan Song. 2024b · 2024
Closest in time.
Bayesian Parameter-Efficient Fine-Tuning for Overcoming Catastrophic Forgetting
Haolin Chen and Philip N Garner. 2024 · 2024
Closest in time.
Take the Bull by the Horns: Hard Sample-Reweighted Continual Training Improves LLM Generalization
Xuxi Chen, Zhendong Wang, Daouda Sow, Junjie Yang, Tianlong Chen, Yingbin Liang, Mingyuan Zhou, and Zhangyang Wang. 2024a · 2024
Closest in time.
Confucius: Iterative tool learning from introspection feedback by easy-to-difficult curriculum. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 18030–18038
Shen Gao, Zhengliang Shi, Minghang Zhu, Bowen Fang, Xin Xin, Pengjie Ren, Zhumin Chen, Jun Ma, and Zhaochun Ren. 2024 · 2024
Closest in time.
Arcee’s MergeKit: A Toolkit for Merging Large Language Models
Charles Goddard, Shamane Siriwardhana, Malikeh Ehghaghi, Luke Meyers, Vlad Karpukhin, Brian Benedict, Mark McQuade, and Jacob Solawetz. 2024 · 2024
Closest in time.
CorpusBrain++: A Continual Generative Pre-Training Framework for Knowledge-Intensive Language Tasks
Jiafeng Guo, Changjiang Zhou, Ruqing Zhang, Jiangui Chen, Maarten de Rijke, Yixing Fan, and Xueqi Cheng. 2024b · 2024
Closest in time.
Q-Tuning: Queue-based prompt tuning for lifelong few-shot language learning
Yanhui Guo, Shaoyuan Xu, Jinmiao Fu, Jia Kevin Liu, Chaosheng Dong, and Bryan Wang. 2024a · 2024
Closest in time.
WilKE: Wise-Layer Knowledge Editor for Lifelong Knowledge Editing
Chenhui Hu, Pengfei Cao, Yubo Chen, Kang Liu, and Jun Zhao. 2024 · 2024
Closest in time.
Mitigating Catastrophic Forgetting in Large Language Models with Self-Synthesized Rehearsal
Jianheng Huang, Leyang Cui, Ante Wang, Chengyi Yang, Xinting Liao, Linfeng Song, Junfeng Yao, and Jinsong Su. 2024a · 2024
Closest in time.
Towards Practical Tool Usage for Continually Learning LLMs
Jerry Huang, Prasanna Parthasarathi, Mehdi Rezagholizadeh, and Sarath Chandar. 2024b · 2024
Closest in time.
Simple and Scalable Strategies to Continually Pre-train Large Language Models
Adam Ibrahim, Benjamin Thérien, Kshitij Gupta, Mats L Richter, Quentin Anthony, Timothée Lesort, Eugene Belilovsky, and Irina Rish. 2024 · 2024
Closest in time.
Trends and Challenges of Real-time Learning in Large Language Models: A Critical Review
Mladjan Jovanovic and Peter Voss. 2024 · 2024
Closest in time.
Examining forgetting in continual pre-training of aligned large language models
Chen-An Li and Hung-Yi Lee. 2024 · 2024
Closest in time.
InfLoRA: Interference-Free Low-Rank Adaptation for Continual Learning
Yan-Shuo Liang and Wu-Jun Li. 2024 · 2024
Closest in time.
Rho-1: Not All Tokens Are What You Need
Zhenghao Lin, Zhibin Gou, Yeyun Gong, Xiao Liu, Yelong Shen, Ruochen Xu, Chen Lin, Yujiu Yang, Jian Jiao, Nan Duan, et al · 2024
Closest in time.
Chameleon: Plug-and-play compositional reasoning with large language models
Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, and Jianfeng Gao. 2024 · 2024
Closest in time.
Making Pre-trained Language Models Better Continual Few-Shot Relation Extractors
Shengkun Ma, Jiale Han, Yi Liang, and Bo Cheng. 2024 · 2024
Closest in time.
HOP to the Next Tasks and Domains for Continual Learning in NLP
Umberto Michieli and Mete Ozay. 2024 · 2024
Closest in time.
Incremental Sequence Labeling: A Tale of Two Shifts
Shengjie Qiu, Junhao Zheng, Zhen Liu, Yicheng Luo, and Qianli Ma. 2024 · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al · 2024
Closest in time.
Analyzing and Reducing Catastrophic Forgetting in Parameter Efficient Tuning
Weijieying Ren, Xinlong Li, Lei Wang, Tianxiang Zhao, and Wei Qin. 2024 · 2024
Closest in time.
Self-generated Replay Memories for Continual Neural Machine Translation
Michele Resta and Davide Bacciu. 2024 · 2024
Closest in time.
Convolutional Prompting meets Language Models for Continual Learning
Anurag Roy, Riddhiman Moulick, Vinay K Verma, Saptarshi Ghosh, and Abir Das. 2024 · 2024
Closest in time.
Continual Learning of Large Language Models: A Comprehensive Survey
Haizhou Shi, Zihao Xu, Hengyi Wang, Weiyi Qin, Wenyuan Wang, Yibin Wang, and Hao Wang. 2024 · 2024
Closest in time.
Continual Learning on Graphs: A Survey
Zonggui Tian, Du Zhang, and Hong-Ning Dai. 2024 · 2024
Closest in time.
Self-Expansion of Pre-trained Models with Mixture of Adapters for Continual Learning
Huiyi Wang, Haodong Lu, Lina Yao, and Dong Gong. 2024c · 2024
Closest in time.
Few-shot Incremental Event Detection
Hao Wang, Hanwen Shi, and Jianyong Duan. 2024e · 2024
Closest in time.
A survey on large language model based autonomous agents
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al · 2024
Closest in time.
A comprehensive survey of continual learning: Theory, method and application
Liyuan Wang, Xingxing Zhang, Hang Su, and Jun Zhu. 2024f · 2024
Closest in time.
Rehearsal-Free Modular and Compositional Continual Learning for Language Models
Mingyang Wang, Heike Adel, Lukas Lange, Jannik Strötgen, and Hinrich Schütze. 2024a · 2024
Closest in time.
Yifan Wang, Yafei Liu, Chufan Shi, Haoling Li, Chen Chen, Haonan Lu, and Yujiu Yang. 2024b · 2024
Closest in time.
Llama pro: Progressive llama with block expansion
Chengyue Wu, Yukang Gan, Yixiao Ge, Zeyu Lu, Jiahao Wang, Ye Feng, Ping Luo, and Ying Shan. 2024a · 2024
Closest in time.
F-MALLOC: Feed-forward Memory Allocation for Continual Learning in Neural Machine Translation
Junhong Wu, Yuchen Liu, and Chengqing Zong. 2024b · 2024
Closest in time.
Continual learning for large language models: A survey
Tongtong Wu, Linhao Luo, Yuan-Fang Li, Shirui Pan, Thuy-Trang Vu, and Gholamreza Haffari. 2024c · 2024
Closest in time.
Less: Selecting influential data for targeted instruction tuning
Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, and Danqi Chen. 2024 · 2024
Closest in time.
Embracing Language Inclusivity and Diversity in CLIP through Continual Language Learning
Bang Yang, Yong Dai, Xuxin Cheng, Yaowei Li, Asif Raza, and Yuexian Zou. 2024b · 2024
Closest in time.
Continual Learning for Smart City: A Survey
Li Yang, Zhipeng Luo, Shiming Zhang, Fei Teng, and Tianrui Li. 2024c · 2024
Closest in time.
Gpt4tools: Teaching large language model to use tools via self-instruction
Rui Yang, Lin Song, Yanwei Li, Sijie Zhao, Yixiao Ge, Xiu Li, and Ying Shan. 2024d · 2024
Closest in time.
MoRAL: MoE Augmented LoRA for LLMs’ Lifelong Learning
Shu Yang, Muhammad Asif Ali, Cheng-Long Wang, Lijie Hu, and Di Wang. 2024a · 2024
Closest in time.
Investigating Continual Pretraining in Large Language Models: Insights and Implications
Çağatay Yıldız, Nishaanth Kanna Ravichandran, Prishruit Punia, Matthias Bethge, and Beyza Ermis. 2024 · 2024
Closest in time.
Continual Few-shot Event Detection via Hierarchical Augmentation Networks
Chenlong Zhang, Pengfei Cao, Yubo Chen, Kang Liu, Zhiqiang Zhang, Mengshu Sun, and Jun Zhao. 2024a · 2024
Closest in time.
COPR: Continual Human Preference Learning via Optimal Policy Regularization
Han Zhang, Lin Gui, Yu Lei, Yuanzhao Zhai, Yehong Zhang, Yulan He, Hui Wang, Yue Yu, Kam-Fai Wong, Bin Liang, et al · 2024
Closest in time.
Set the Clock: Temporal Alignment of Pretrained Language Models
Bowen Zhao, Zander Brumbaugh, Yizhong Wang, Hannaneh Hajishirzi, and Noah A Smith. 2024 · 2024
Closest in time.
Beyond Anti-Forgetting: Multimodal Continual Instruction Tuning with Positive Forward Transfer
Junhao Zheng, Qianli Ma, Zhen Liu, Binquan Wu, and Huawen Feng. 2024a · 2024
Closest in time.
Concept-1K: A Novel Benchmark for Instance Incremental Learning
Junhao Zheng, Shengjie Qiu, and Qianli Ma. 2024b · 2024
Closest in time.
Balancing the Causal Effects in Class-Incremental Learning
Junhao Zheng, Ruiyan Wang, Chongzhi Zhang, Huawen Feng, and Qianli Ma. 2024c · 2024
Closest in time.
Continual Learning with Pre-Trained Models: A Survey
Da-Wei Zhou, Hai-Long Sun, Jingyi Ning, Han-Jia Ye, and De-Chuan Zhan. 2024 · 2024
Closest in time.
Overcoming Catastrophic Forgetting beyond Continual Learning: Balanced Training for Neural Machine Translation. In Proceedings of Annual Meeting of the Association for Computational Linguistics . 2023–2036
Chenze Shao and Yang Feng. 2022 · 2036
Closest in time.