Fetching the paper…
Reading the bibliography…
The recent success of large language models (LLMs) trained on static, pre-collected, general datasets has sparked numerous research directions and applications.
Catastrophic interference in connectionist networks: The sequential learning problem
M. McCloskey and N. J. Cohen · 1989
Earlier work this paper cites.
Why there are complementary learning systems in the hippocampus and neocortex: insights from the successes and failures of connectionist models of learning and memory
J. L. McClelland, B. L. McNaughton, and R. C. O’Reilly · 1995
Earlier work this paper cites.
Principles of neural science
E. R. Kandel, J. H. Schwartz, T. M. Jessell, S. Siegelbaum, A. J. Hudspeth, S. Mack, et al · 2000
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Brain imaging of language plasticity in adopted adults: Can a second language replace the first?
C. Pallier, S. Dehaene, J.-B. Poline, D. LeBihan, A.-M. Argenti, E. Dupoux, and J. Mehler · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
S. Banerjee and A. Lavie · 2005
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
A theory of learning from different domains
S. Ben-David, J. Blitzer, K. Crammer, A. Kulesza, F. Pereira, and J. W. Vaughan · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
A. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts · 2011
Earlier work this paper cites.
Findings of the 2014 workshop on statistical machine translation
O. Bojar, C. Buck, C. Federmann, B. Haddow, P. Koehn, J. Leveling, C. Monz, P. Pecina, M. Post, H. Saint-Amand, et al · 2014
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling, 2014
C. Chelba, T. Mikolov, M. Schuster, Q. Ge, T. Brants, P. Koehn, and T. Robinson · 2014
Earlier work this paper cites.
ReferItGame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg · 2014
Earlier work this paper cites.
Character-level convolutional networks for text classification
X. Zhang, J. Zhao, and Y. LeCun · 2015
Earlier work this paper cites.
Dbpedia abstracts: A large-scale, open, multilingual nlp training corpus
M. Brümmer, M. Dojchinovski, and S. Hellmann · 2016
Earlier work this paper cites.
Domain-adversarial training of neural networks
Y. Ganin, E. Ustinova, H. Ajakan, P. Germain, H. Larochelle, F. Laviolette, M. Marchand, and V. Lempitsky · 2016
Earlier work this paper cites.
Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering
R. He and J. McAuley · 2016
Earlier work this paper cites.
Wikireading: A novel large-scale language understanding task over wikipedia
D. Hewlett, A. Lacoste, L. Jones, I. Polosukhin, A. Fandrianto, J. Han, M. Kelcey, and D. Berthelot · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions, 2016
J. Mao, J. Huang, A. Toshev, O. Camburu, A. Yuille, and K. Murphy · 2016
Earlier work this paper cites.
Theoretical foundations of multi-task lifelong learning
A. Pentina · 2016
Earlier work this paper cites.
A. A. Rusu, N. C. Rabinowitz, G. Desjardins, H. Soyer, J. Kirkpatrick, K. Kavukcuoglu, R. Pascanu, and R. Hadsell · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering, 2017
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
Learning what is essential in questions
D. Khashabi, T. Khot, A. Sabharwal, and D. Roth · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, et al · 2017
Earlier work this paper cites.
Race: Large-scale reading comprehension dataset from examinations
G. Lai, Q. Xie, H. Liu, Y. Yang, and E. Hovy · 2017
Earlier work this paper cites.
Zero-shot relation extraction via reading comprehension
O. Levy, M. Seo, E. Choi, and L. Zettlemoyer · 2017
Earlier work this paper cites.
Learning without forgetting
Z. Li and D. Hoiem · 2017
Earlier work this paper cites.
Gradient episodic memory for continual learning
D. Lopez-Paz and M. Ranzato · 2017
Earlier work this paper cites.
icarl: Incremental classifier and representation learning
S.-A. Rebuffi, A. Kolesnikov, G. Sperl, and C. H. Lampert · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean · 2017
Earlier work this paper cites.
Tl; dr: Mining reddit to learn automatic summarization
M. Völske, M. Potthast, S. Syed, and B. Stein · 2017
Earlier work this paper cites.
Memory aware synapses: Learning what (not) to forget
R. Aljundi, F. Babiloni, M. Elhoseiny, M. Rohrbach, and T. Tuytelaars · 2018
Earlier work this paper cites.
Caselaw access project, 2018
Caselaw Access Project · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
T-rex: A large scale alignment of natural language with knowledge base triples
H. Elsahar, P. Vougiouklis, A. Remaci, C. Gravier, J. Hare, F. Laforest, and E. Simperl · 2018
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people, 2018
D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham · 2018
Earlier work this paper cites.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
D. Khashabi, S. Chaturvedi, M. Roth, S. Upadhyay, and D. Roth · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for squad
P. Rajpurkar, R. Jia, and P. Liang · 2018
Earlier work this paper cites.
Learning to learn without forgetting by maximizing transfer and minimizing interference
M. Riemer, I. Cases, R. Ajemian, M. Liu, I. Rish, Y. Tu, and G. Tesauro · 2018
Earlier work this paper cites.
Online structured laplace approximations for overcoming catastrophic forgetting
H. Ritter, A. Botev, and D. Barber · 2018
Earlier work this paper cites.
Progress & compress: A scalable framework for continual learning
J. Schwarz, W. Czarnecki, J. Luketina, A. Grabska-Barwinska, Y. W. Teh, R. Pascanu, and R. Hadsell · 2018
Earlier work this paper cites.
Fever: a large-scale dataset for fact extraction and verification
J. Thorne, A. Vlachos, C. Christodoulopoulos, and A. Mittal · 2018
Earlier work this paper cites.
The kanerva machine: A generative distributed memory
Y. Wu, G. Wayne, A. Graves, and T. Lillicrap · 2018
Earlier work this paper cites.
Finbert: Financial sentiment analysis with pre-trained language models, 2019
D. Araci · 2019
Earlier work this paper cites.
Efficient lifelong learning with a-gem
A. Chaudhry, M. Ranzato, M. Rohrbach, and M. Elhoseiny · 2019
Earlier work this paper cites.
On tiny episodic memories in continual learning
A. Chaudhry, M. Rohrbach, M. Elhoseiny, T. Ajanthan, P. K. Dokania, P. H. Torr, and M. Ranzato · 2019
Earlier work this paper cites.
Quoref: A reading comprehension dataset with questions requiring coreferential reasoning
P. Dasigi, N. F. Liu, A. Marasović, N. A. Smith, and M. Gardner · 2019
Earlier work this paper cites.
Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs
D. Dua, Y. Wang, P. Dasigi, G. Stanovsky, S. Singh, and M. Gardner · 2019
Earlier work this paper cites.
Uncertainty-guided continual learning with bayesian neural networks
S. Ebrahimi, M. Elhoseiny, T. Darrell, and M. Rohrbach · 2019
Earlier work this paper cites.
Openwebtext corpus, 2019
A. Gokaslan and V. Cohen · 2019
Earlier work this paper cites.
Cosmos qa: Machine reading comprehension with contextual commonsense reasoning, 2019
L. Huang, R. L. Bras, C. Bhagavatula, and Y. Choi · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering, 2019
D. A. Hudson and C. D. Manning · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, et al · 2019
Earlier work this paper cites.
Reasoning over paragraph effects in situations, 2019
K. Lin, O. Tafjord, P. Clark, and M. Gardner · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Earlier work this paper cites.
Ocr-vqa: Visual question answering by reading text in images
A. Mishra, S. Shekhar, A. K. Singh, and A. Chakraborty · 2019
Earlier work this paper cites.
Justifying recommendations using distantly-labeled reviews and fine-grained aspects
J. Ni, J. Li, and J. McAuley · 2019
Earlier work this paper cites.
Language models as knowledge bases?
F. Petroni, T. Rocktäschel, S. Riedel, P. Lewis, A. Bakhtin, Y. Wu, and A. Miller · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale, 2019
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi · 2019
Earlier work this paper cites.
Towards vqa models that can read, 2019
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach · 2019
Earlier work this paper cites.
Large scale incremental learning
Y. Wu, Y. Chen, L. Wang, Y. Ye, Z. Liu, Y. Guo, and Y. Fu · 2019
Earlier work this paper cites.
Bert post-training for review reading comprehension and aspect-based sentiment analysis, 2019
H. Xu, B. Liu, L. Shu, and P. S. Yu · 2019
Earlier work this paper cites.
Defending against neural fake news
R. Zellers, A. Holtzman, H. Rashkin, Y. Bisk, A. Farhadi, F. Roesner, and Y. Choi · 2019
Earlier work this paper cites.
"going on a vacation" takes longer than "going for a walk": A study of temporal commonsense understanding, 2019
B. Zhou, D. Khashabi, Q. Ning, and D. Roth · 2019
Earlier work this paper cites.
The pushshift reddit dataset, 2020
J. Baumgartner, S. Zannettou, B. Keegan, M. Squire, and J. Blackburn · 2020
Earlier work this paper cites.
Continual lifelong learning in natural language processing: A survey
M. Biesialska, K. Biesialska, and M. R. Costa-jussà · 2020
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Y. Bisk, R. Zellers, J. Gao, Y. Choi, et al · 2020
Earlier work this paper cites.
Machine unlearning, 2020
L. Bourtoule, V. Chandrasekaran, C. A. Choquette-Choo, H. Jia, A. Travers, B. Zhang, D. Lie, and N. Papernot · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Dark experience for general continual learning: a strong, simple baseline
P. Buzzega, M. Boschini, A. Porrello, D. Abati, and S. Calderara · 2020
Earlier work this paper cites.
Recall and learn: Fine-tuning deep pretrained language models with less forgetting
S. Chen, Y. Hou, Y. Cui, W. Che, T. Liu, and X. Yu · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Earlier work this paper cites.
Adversarial continual learning
S. Ebrahimi, F. Meier, R. Calandra, T. Darrell, and M. Rohrbach · 2020
Earlier work this paper cites.
Don’t stop pretraining: Adapt language models to domains and tasks
S. Gururangan, A. Marasović, S. Swayamdipta, K. Lo, I. Beltagy, D. Downey, and N. A. Smith · 2020
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Earlier work this paper cites.
Qasc: A dataset for question answering via sentence composition, 2020
T. Khot, P. Clark, M. Guerquin, P. Jansen, and A. Sabharwal · 2020
Earlier work this paper cites.
S2ORC: The semantic scholar open research corpus
K. Lo, L. L. Wang, M. Neumann, R. Kinney, and D. Weld · 2020
Earlier work this paper cites.
Rehearsal-free continual learning over small non-i.i.d. batches, 2020
V. Lomonaco, D. Maltoni, and L. Pellegrini · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
A. Sinitsin, V. Plokhotnyuk, D. Pyrkin, S. Popov, and A. Babenko · 2020
Earlier work this paper cites.
Ernie 2.0: A continual pre-training framework for language understanding
Y. Sun, S. Wang, Y. Li, S. Feng, H. Tian, H. Wu, and H. Wang · 2020
Earlier work this paper cites.
Dynamic language models for continuously evolving content
S. Amba Hombaiah, T. Chen, M. Zhang, M. Bendersky, and M. Najork · 2021
Earlier work this paper cites.
Editing factual knowledge in language models
N. De Cao, W. Aziz, and I. Titov · 2021
Earlier work this paper cites.
Domain-specific language model pretraining for biomedical natural language processing
Y. Gu, R. Tinn, H. Cheng, M. Lucas, N. Usuyama, X. Liu, T. Naumann, J. Gao, and H. Poon · 2021
Earlier work this paper cites.
ECONET: Effective continual pretraining of language models for event temporal reasoning
R. Han, X. Ren, and N. Peng · 2021
Earlier work this paper cites.
Do language models have beliefs? methods for detecting, updating, and visualizing model beliefs
P. Hase, M. Diab, A. Celikyilmaz, X. Li, Z. Kozareva, V. Stoyanov, M. Bansal, and S. Iyer · 2021
Earlier work this paper cites.
Analyzing the forgetting problem in pretrain-finetuning of open-domain dialogue response models
T. He, J. Liu, K. Cho, M. Ott, B. Liu, J. Glass, and F. Peng · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Earlier work this paper cites.
Achieving forgetting prevention and knowledge transfer in continual learning
Z. Ke, B. Liu, N. Ma, H. Xu, and S. Lei · 2021
Earlier work this paper cites.
Mind the gap: Assessing temporal generalization in neural language models
A. Lazaridou, A. Kuncoro, E. Gribovskaya, D. Agrawal, A. Liska, T. Terzi, M. Gimenez, C. de Masson d’Autume, T. Kocisky, S. Ruder, et al · 2021
Earlier work this paper cites.
Codexglue: A machine learning benchmark dataset for code understanding and generation, 2021
S. Lu, D. Guo, S. Ren, J. Huang, A. Svyatkovskiy, A. Blanco, C. Clement, D. Drain, D. Jiang, D. Tang, G. Li, L. Zhou, L. Shou, L. Zhou, M. Tufano, M. Gong, M. Zhou, N. Duan, N. Sundaresan, S. K. Deng, S. Fu, and S. Liu · 2021
Earlier work this paper cites.
D. McCaffary · 2021
Earlier work this paper cites.
Natural instructions: Benchmarking generalization to new tasks from natural language instructions
S. Mishra, D. Khashabi, C. Baral, and H. Hajishirzi · 2021
Earlier work this paper cites.
E. Mitchell, C. Lin, A. Bosselut, C. Finn, and C. D. Manning · 2021
Earlier work this paper cites.
Revisiting catastrophic forgetting in class incremental learning
Z. Ni, H. Shi, S. Tang, L. Wei, Q. Tian, and Y. Zhuang · 2021
Earlier work this paper cites.
Lfpt5: A unified framework for lifelong few-shot language learning based on prompt tuning of t5
C. Qin and S. Joty · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Model zoo: A growing" brain" that learns continually
R. Ramesh and P. Chaudhari · 2021
Cited alongside, same era.
Hopfield networks is all you need, 2021
H. Ramsauer, B. Schäfl, J. Lehner, P. Seidl, M. Widrich, T. Adler, L. Gruber, M. Holzleitner, M. Pavlović, G. K. Sandve, V. Greiff, D. Kreil, M. Kopp, G. Klambauer, J. Brandstetter, and S. Hochreiter · 2021
Cited alongside, same era.
Continual domain-tuning for pretrained language models, 2021
S. Rongali, A. Jagannatha, B. P. S. Rawat, and H. Yu · 2021
Cited alongside, same era.
Get your vitamin c! robust fact verification with contrastive evidence
T. Schuster, A. Fisch, and R. Barzilay · 2021
Cited alongside, same era.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
B. Wang and A. Komatsuzaki · 2021
Cited alongside, same era.
Computationally budgeted continual learning: What does matter?
A. Prabhu, H. A. Al Kader Hammoud, P. K. Dokania, P. H. Torr, S.-N. Lim, B. Ghanem, and A. Bibi · 2023
Later among the works it cites.
Online continual learning without the storage constraint, 2023
A. Prabhu, Z. Cai, P. Dokania, P. Torr, V. Koltun, and O. Sener · 2023
Later among the works it cites.
Recyclable tuning for continual pre-training
Y. Qin, C. Qian, X. Han, Y. Lin, H. Wang, R. Xie, Z. Liu, M. Sun, and J. Zhou · 2023
Later among the works it cites.
A. N. Rubungo, C. Arnold, B. P. Rand, and A. B. Dieng · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K-Adapter: Infusing Knowledge into Pre-Trained Models with Adapters
R. Wang, D. Tang, N. Duan, Z. Wei, X. Huang, J. Ji, G. Cao, D. Jiang, and M. Zhou · 2021
Cited alongside, same era.
Codet5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation
Y. Wang, W. Wang, S. Joty, and S. C. Hoi · 2021
Cited alongside, same era.
Finetuned language models are zero-shot learners
J. Wei, M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le · 2021
Cited alongside, same era.
Pretrained language model in continual learning: A comparative study
T. Wu, M. Caccia, Z. Li, Y.-F. Li, G. Qi, and G. Haffari · 2021
Cited alongside, same era.
Pre-training text-to-text transformers for concept-centric common sense
W. Zhou, D.-H. Lee, R. K. Selvam, S. Lee, B. Y. Lin, and X. Ren · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Cited alongside, same era.
Fairlex: A multilingual benchmark for evaluating fairness in legal text processing
I. Chalkidis, T. Pasini, S. Zhang, L. Tomada, S. F. Schwemer, and A. Søgaard · 2022
Cited alongside, same era.
F. Sarfraz, E. Arani, and B. Zonooz · 2023
Later among the works it cites.
Trillion dollar words: A new financial dataset, task & market analysis, 2023
A. Shah, S. Paturi, and S. Chava · 2023
Later among the works it cites.
Class-incremental learning based on label generation
Y. Shao, Y. Guo, D. Zhao, and B. Liu · 2023
Later among the works it cites.
Conpet: Continual parameter-efficient tuning for large language models, 2023
C. Song, X. Han, Z. Zeng, K. Li, C. Chen, Z. Liu, M. Sun, and T. Yang · 2023
Later among the works it cites.
Efficient continue training of temporal language model with structural information
Z. Su, J. Li, Z. Zhang, Z. Zhou, and M. Zhang · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Later among the works it cites.
Starcoder: may the source be with you!, 2023
S. Team · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Later among the works it cites.
Orthogonal subspace learning for language model continual learning
X. Wang, T. Chen, Q. Ge, H. Xia, R. Bao, R. Zheng, Q. Zhang, T. Gui, and X. Huang · 2023
Later among the works it cites.
Trace: A comprehensive benchmark for continual learning in large language models, 2023
X. Wang, Y. Zhang, T. Chen, S. Gao, S. Jin, X. Yang, Z. Xi, R. Zheng, Y. Zou, T. Gui, Q. Zhang, and X. Huang · 2023
Later among the works it cites.
Codet5+: Open code large language models for code understanding and generation, 2023
Y. Wang, H. Le, A. D. Gotmare, N. D. Q. Bui, J. Li, and S. C. H. Hoi · 2023
Later among the works it cites.
On the usage of continual learning for out-of-distribution generalization in pre-trained language models of code
M. Weyssow, X. Zhou, K. Kim, D. Lo, and H. Sahraoui · 2023
Later among the works it cites.
Overcoming catastrophic forgetting in massively multilingual continual learning
G. Winata, L. Xie, K. Radhakrishnan, S. Wu, X. Jin, P. Cheng, M. Kulkarni, and D. Preotiuc-Pietro · 2023
Later among the works it cites.
Continual learning with low rank adaptation
M. Wistuba, P. T. Sivaprasad, L. Balles, and G. Zappella · 2023
Later among the works it cites.
Pmc-llama: Towards building open-source language models for medicine
C. Wu, W. Lin, X. Zhang, Y. Zhang, Y. Wang, and W. Xie · 2023
Later among the works it cites.
Bloomberggpt: A large language model for finance
S. Wu, O. Irsoy, S. Lu, V. Dabravolski, M. Dredze, S. Gehrmann, P. Kambadur, D. S. Rosenberg, and G. Mann · 2023
Later among the works it cites.
Quert: Continual pre-training of language model for query understanding in travel domain search
J. Xie, Y. Liang, J. Liu, Y. Xiao, B. Wu, and S. Ni · 2023
Later among the works it cites.
PIXIU: A large language model, instruction data and evaluation benchmark for finance
Q. Xie, W. Han, X. Zhang, Y. Lai, M. Peng, A. Lopez-Lira, and J. Huang · 2023
Later among the works it cites.
Efficient continual pre-training for building domain specific large language models, 2023
Y. Xie, K. Aggarwal, and A. Ahmad · 2023
Later among the works it cites.
Wizardlm: Empowering large language models to follow complex instructions
C. Xu, Q. Sun, K. Zheng, X. Geng, P. Zhao, J. Feng, C. Tao, and D. Jiang · 2023
Later among the works it cites.
S. Xue, F. Zhou, Y. Xu, H. Zhao, S. Xie, Q. Dai, C. Jiang, J. Zhang, J. Zhou, D. Xiu, and H. Mei · 2023
Later among the works it cites.
Af adapter: Continual pretraining for building chinese biomedical language model
Y. Yan, K. Xue, X. Shi, Q. Ye, J. Liu, and T. Ruan · 2023
Later among the works it cites.
Melo: Enhancing model editing with neuron-indexed dynamic lora
L. Yu, Q. Chen, J. Zhou, and L. He · 2023
Later among the works it cites.
Mammoth: Building math generalist models through hybrid instruction tuning
X. Yue, X. Qu, G. Zhang, Y. Fu, W. Huang, H. Sun, Y. Su, and W. Chen · 2023
Later among the works it cites.
Investigating the catastrophic forgetting in multimodal large language models, 2023
Y. Zhai, S. Tong, X. Li, M. Cai, Q. Qu, Y. J. Lee, and Y. Ma · 2023
Later among the works it cites.
Copf: Continual learning human preference through optimal policy fitting
H. Zhang, L. Gui, Y. Zhai, H. Wang, Y. Lei, and R. Xu · 2023
Later among the works it cites.
Xuanyuan 2.0: A large chinese financial chat model with hundreds of billions parameters
X. Zhang and Q. Yang · 2023
Later among the works it cites.
CITB: A benchmark for continual instruction tuning
Z. Zhang, M. Fang, L. Chen, and M.-R. Namazi-Rad · 2023
Later among the works it cites.
C-STANCE: A large dataset for Chinese zero-shot stance detection
C. Zhao, Y. Li, and C. Caragea · 2023
Later among the works it cites.
Learn or recall? revisiting incremental learning with pre-trained language models, 2023
J. Zheng, S. Qiu, and Q. Ma · 2023
Later among the works it cites.
Preventing zero-shot transfer degradation in continual learning of vision-language models
Z. Zheng, M. Ma, K. Wang, Z. Qin, X. Yue, and Y. You · 2023
Later among the works it cites.
Marinegpt: Unlocking secrets of ocean to the public
Z. Zheng, J. Zhang, T. Vu, S. Diao, Y. H. W. Tim, and S. Yeung · 2023
Later among the works it cites.
Hippocrates: An open-source framework for advancing large language models in healthcare
E. C. Acikgoz, O. B. İnce, R. Bench, A. A. Boz, İ. Kesen, A. Erdem, and E. Erdem · 2024
Closest in time.
Structured code representations enable data-efficient adaptation of code language models, 2024
M. Agarwal, Y. Shen, B. Wang, Y. Kim, and J. Chen · 2024
Closest in time.
Transformers for supervised online continual learning, 2024
J. Bornschein, Y. Li, and A. Rannen-Triki · 2024
Closest in time.
Generative multi-modal models are good class incremental learners
X. Cao, H. Lu, L. Huang, X. Liu, and M.-M. Cheng · 2024
Closest in time.
Coin: A benchmark of continual instruction tuning for multimodel large language model, 2024
C. Chen, J. Zhu, X. Luo, H. Shen, L. Gao, and J. Song · 2024
Closest in time.
Take the bull by the horns: Hard sample-reweighted continual training improves llm generalization
X. Chen, Z. Wang, D. Sow, J. Yang, T. Chen, Y. Liang, M. Zhou, and Z. Wang · 2024
Closest in time.
Parameterizing context: Unleashing the power of parameter-efficient fine-tuning and in-context tuning for continual table semantic parsing
Y. Chen, S. Zhang, G. Qi, and X. Guo · 2024
Closest in time.
Adapting large language models via reading comprehension, 2024
D. Cheng, S. Huang, and F. Wei · 2024
Closest in time.
Saullm-7b: A pioneering large language model for law, 2024
P. Colombo, T. P. Pires, M. Boudiaf, D. Culver, R. Melo, C. Corro, A. F. T. Martins, F. Esposito, V. L. Raposo, S. Morgado, and M. Desa · 2024
Closest in time.
Larimar: Large language models with episodic memory control
P. Das, S. Chaudhury, E. Nelson, I. Melnyk, S. Swaminathan, S. Dai, A. Lozano, G. Kollias, V. Chenthamarakshan, S. Dan, et al · 2024
Closest in time.
Sailor: Open language models for south-east asia
L. Dou, Q. Liu, G. Zeng, J. Guo, J. Zhou, W. Lu, and M. Lin · 2024
Closest in time.
Continual pre-training for cross-lingual llm adaptation: Enhancing japanese language capabilities
K. Fujii, T. Nakamura, M. Loem, H. Iida, M. Ohi, K. Hattori, H. Shota, S. Mizuki, R. Yokota, and N. Okazaki · 2024
Closest in time.
Tic-clip: Continual training of clip models
S. Garg, M. Farajtabar, H. Pouransari, R. Vemulapalli, S. Mehta, O. Tuzel, V. Shankar, and F. Faghri · 2024
Closest in time.
Continual learning under language shift, 2024
E. Gogoulou, T. Lesort, M. Boman, and J. Nivre · 2024
Closest in time.
Deepseek-coder: When the large language model meets programming – the rise of code intelligence, 2024
D. Guo, Q. Zhu, D. Yang, Z. Xie, K. Dong, W. Zhang, G. Chen, X. Bi, Y. Wu, Y. K. Li, F. Luo, Y. Xiong, and W. Liang · 2024
Closest in time.
Don’t half-listen: Capturing key-part information in continual instruction tuning, 2024
Y. He, X. Huang, M. Tang, L. Meng, X. Li, W. Lin, W. Zhang, and Y. Gao · 2024
Closest in time.
Wilke: Wise-layer knowledge editor for lifelong knowledge editing, 2024
C. Hu, P. Cao, Y. Chen, K. Liu, and J. Zhao · 2024
Closest in time.
Mitigating catastrophic forgetting in large language models with self-synthesized rehearsal, 2024
J. Huang, L. Cui, A. Wang, C. Yang, X. Liao, L. Song, J. Yao, and J. Su · 2024
Closest in time.
Simple and scalable strategies to continually pre-train large language models
A. Ibrahim, B. Thérien, K. Gupta, M. L. Richter, Q. Anthony, T. Lesort, E. Belilovsky, and I. Rish · 2024
Closest in time.
Ai alignment: A comprehensive survey, 2024
J. Ji, T. Qiu, B. Chen, B. Zhang, H. Lou, K. Wang, Y. Duan, Z. He, J. Zhou, Z. Zhang, F. Zeng, K. Y. Ng, J. Dai, X. Pan, A. O’Gara, Y. Lei, H. Xu, B. Tse, J. Fu, S. McAleer, Y. Yang, Y. Wang, S.-C. Zhu, Y. Guo, and W. Gao · 2024
Closest in time.
Instruction-tuned language models are better knowledge learners, 2024
Z. Jiang, Z. Sun, W. Shi, P. Rodriguez, C. Zhou, G. Neubig, X. V. Lin, W. tau Yih, and S. Iyer · 2024
Closest in time.
What will my model forget? forecasting forgotten examples in language model refinement, 2024
X. Jin and X. Ren · 2024
Closest in time.
Examining forgetting in continual pre-training of aligned large language models, 2024
C.-A. Li and H.-Y. Lee · 2024
Closest in time.
Blade: Enhancing black-box large language models with small domain-specific models
H. Li, Q. Ai, J. Chen, Q. Dong, Z. Wu, Y. Liu, C. Chen, and Q. Tian · 2024
Closest in time.
Mitigating the alignment tax of rlhf, 2024
Y. Lin, H. Lin, W. Xiong, S. Diao, J. Liu, J. Zhang, R. Pan, H. Wang, W. Hu, H. Zhang, H. Dong, R. Pi, H. Zhao, N. Jiang, H. Ji, Y. Yao, and T. Zhang · 2024
Closest in time.
Rho-1: Not all tokens are what you need
Z. Lin, Z. Gou, Y. Gong, X. Liu, Y. Shen, R. Xu, C. Lin, Y. Yang, J. Jiao, N. Duan, et al · 2024
Closest in time.
T. Nakamura, M. Mishra, S. Tedeschi, Y. Chai, J. T. Stillerman, F. Friedrich, P. Yadav, T. Laud, V. M. Chien, T. Y. Zhuo, et al · 2024
Closest in time.
Ircoder: Intermediate representations make language models robust multilingual code generators, 2024
I. Paul, J. Luo, G. Glavaš, and I. Gurevych · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
M. Reid, N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J.-b. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser, et al · 2024
Closest in time.
Code llama: Open foundation models for code, 2024
B. Rozière, J. Gehring, F. Gloeckle, S. Sootla, I. Gat, X. E. Tan, Y. Adi, J. Liu, R. Sauvestre, T. Remez, J. Rapin, A. Kozhevnikov, I. Evtimov, J. Bitton, M. Bhatt, C. C. Ferrer, A. Grattafiori, W. Xiong, A. Défossez, J. Copet, F. Azhar, H. Touvron, L. Martin, N. Usunier, T. Scialom, and G. Synnaeve · 2024
Closest in time.
Tag-llm: Repurposing general-purpose llms for specialized domains
J. Shen, N. Tenenholtz, J. B. Hall, D. Alvarez-Melis, and N. Fusi · 2024
Closest in time.
A unified approach to domain incremental learning with memory: Theory and algorithm
H. Shi and H. Wang · 2024
Closest in time.
Dolma: an open corpus of three trillion tokens for language model pretraining research, 2024
L. Soldaini, R. Kinney, A. Bhagia, D. Schwenk, D. Atkinson, R. Authur, B. Bogin, K. Chandu, J. Dumas, Y. Elazar, V. Hofmann, A. H. Jha, S. Kumar, L. Lucy, X. Lyu, N. Lambert, I. Magnusson, J. Morrison, N. Muennighoff, A. Naik, C. Nam, M. E. Peters, A. Ravichander, K. Richardson, Z. Shen, E. Strubell, N. Subramani, O. Tafjord, P. Walsh, L. Zettlemoyer, N. A. Smith, H. Hajishirzi, I. Beltagy, D. Groeneveld, J. Dodge, and K. Lo · 2024
Closest in time.
Code needs comments: Enhancing code llms with comment augmentation, 2024
D. Song, H. Guo, Y. Zhou, S. Xing, Y. Wang, Z. Song, W. Zhang, Q. Guo, H. Yan, X. Qiu, and D. Lin · 2024
Closest in time.
A survey of neural code intelligence: Paradigms, advances and beyond, 2024
Q. Sun, Z. Chen, F. Xu, K. Cheng, C. Ma, Z. Yin, J. Wang, C. Han, R. Zhu, S. Yuan, Q. Guo, X. Qiu, P. Yin, X. Li, F. Yuan, L. Kong, X. Li, and Z. Wu · 2024
Closest in time.
K. Takahashi, T. Omi, K. Arima, and T. Ishigaki · 2024
Closest in time.
Deepseek llm: Scaling open-source language models with longtermism, 2024
D.-A. Team · 2024
Closest in time.
Starcoder 2 and the stack v2: The next generation, 2024
S. Team · 2024
Closest in time.
Climategpt: Towards ai synthesizing interdisciplinary research on climate change
D. Thulke, Y. Gao, P. Pelser, R. Brune, R. Jalota, F. Fok, M. Ramos, I. van Wyk, A. Nasir, H. Goldstein, et al · 2024
Closest in time.
Continual learning: Applications and the road forward, 2024
E. Verwimp, R. Aljundi, S. Ben-David, M. Bethge, A. Cossu, A. Gepperth, T. L. Hayes, E. Hüllermeier, C. Kanan, D. Kudithipudi, C. H. Lampert, M. Mundt, R. Pascanu, A. Popescu, A. S. Tolias, J. van de Weijer, B. Liu, V. Lomonaco, T. Tuytelaars, and G. M. van de Ven · 2024
Closest in time.
A comprehensive survey of continual learning: Theory, method and application
L. Wang, X. Zhang, H. Su, and J. Zhu · 2024
Closest in time.
Wise: Rethinking the knowledge memory for lifelong model editing of large language models
P. Wang, Z. Li, N. Zhang, Z. Xu, Y. Yao, Y. Jiang, P. Xie, F. Huang, and H. Chen · 2024
Closest in time.
Inscl: A data-efficient continual learning paradigm for fine-tuning large language models with instructions, 2024
Y. Wang, Y. Liu, C. Shi, H. Li, C. Chen, H. Lu, and Y. Yang · 2024
Closest in time.
Codeclm: Aligning language models with tailored synthetic data
Z. Wang, C.-L. Li, V. Perot, L. T. Le, J. Miao, Z. Zhang, C.-Y. Lee, and T. Pfister · 2024
Closest in time.
Llama pro: Progressive llama with block expansion, 2024
C. Wu, Y. Gan, Y. Ge, Z. Lu, J. Wang, Y. Feng, P. Luo, and Y. Shan · 2024
Closest in time.
Continual learning for large language models: A survey, 2024
T. Wu, L. Luo, Y.-F. Li, S. Pan, T.-T. Vu, and G. Haffari · 2024
Closest in time.
Me llama: Foundation large language models for medical applications
Q. Xie, Q. Chen, A. Chen, C. Peng, Y. Hu, F. Lin, X. Peng, J. Huang, J. Zhang, V. Keloth, et al · 2024
Closest in time.
Data selection for language models via importance resampling
S. M. Xie, S. Santurkar, T. Ma, and P. S. Liang · 2024
Closest in time.
Moral: Moe augmented lora for llms’ lifelong learning, 2024
S. Yang, M. A. Ali, C.-L. Wang, L. Hu, and D. Wang · 2024
Closest in time.
Pllama: An open-source large language model for plant science
X. Yang, J. Gao, W. Xue, and E. Alexandersson · 2024
Closest in time.
Reawakening knowledge: Anticipatory recovery from catastrophic interference via structured training, 2024
Y. Yang, M. Jones, M. C. Mozer, and M. Ren · 2024
Closest in time.
Recent advances of foundation language models-based continual learning: A survey
Y. Yang, J. Zhou, X. Ding, T. Huai, S. Liu, Q. Chen, L. He, and Y. Xie · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao, and K. Narasimhan · 2024
Closest in time.
Investigating continual pretraining in large language models: Insights and implications
Ç. Yıldız, N. K. Ravichandran, P. Punia, M. Bethge, and B. Ermis · 2024
Closest in time.
Sciglm: Training scientific language models with self-reflective instruction annotation and tuning
D. Zhang, Z. Hu, S. Zhoubian, Z. Du, K. Yang, Z. Wang, Y. Yue, Y. Dong, and J. Tang · 2024
Closest in time.
Instruction tuning for large language models: A survey, 2024
S. Zhang, L. Dong, X. Li, S. Zhang, X. Sun, S. Wang, J. Li, R. Hu, T. Zhang, F. Wu, and G. Wang · 2024
Closest in time.
Large language model can continue evolving from mistakes
H. Zhao, H. Han, J. Shi, C. Du, J. Liang, and Y. Xiao · 2024
Closest in time.
Reconstruct before query: Continual missing modality learning with decomposed prompt collaboration, 2024
S. Zhao, X. Zou, T. Yu, and H. Xu · 2024
Closest in time.
Sapt: A shared attention framework for parameter-efficient continual learning of large language models, 2024
W. Zhao, S. Wang, Y. Hu, Y. Zhao, B. Qin, X. Zhang, Q. Yang, D. Xu, and W. Che · 2024
Closest in time.
Beyond anti-forgetting: Multimodal continual instruction tuning with positive forward transfer, 2024
J. Zheng, Q. Ma, Z. Liu, B. Wu, and H. Feng · 2024
Closest in time.
Model tailor: Mitigating catastrophic forgetting in multi-modal large language models, 2024
D. Zhu, Z. Sun, Z. Li, T. Shen, K. Yan, S. Ding, K. Kuang, and C. Wu · 2024
Closest in time.