Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are routinely pre-trained on billions of tokens, only to restart the process over again once new data becomes available.
Orthogonal gradient descent for continual learning
Farajtabar, M., Azizan, N., Mott, A., and Li, A · 1910
Earlier work this paper cites.
Continual unsupervised representation learning
Rao, D., Visin, F., Rusu, A. A., Teh, Y. W., Pascanu, R., and Hadsell, R · 1910
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
French, R. M · 1999
Earlier work this paper cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Gururangan, S., Marasović, A., Swayamdipta, S., Lo, K., Beltagy, I., Downey, D., and Smith, N. A · 2004
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2005
Earlier work this paper cites.
The impact of non-stationarity on generalisation in deep reinforcement learning
Igl, M., Farquhar, G., Luketina, J., Boehmer, W., and Whiteson, S · 2006
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al · 2017
Earlier work this paper cites.
Automatic Differentiation in PyTorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Earlier work this paper cites.
icarl: Incremental classifier and representation learning
Rebuffi, S.-A., Kolesnikov, A., Sperl, G., and Lampert, C. H · 2017
Earlier work this paper cites.
Continual learning in generative adversarial nets
Seff, A., Beatson, A., Suo, D., and Liu, H · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2018
Earlier work this paper cites.
Variational continual learning
Nguyen, C. V., Li, Y., Bui, T. D., and Turner, R. E · 2018
Earlier work this paper cites.
IMHO fine-tuning improves claim detection
Chakrabarty, T., Hidey, C., and McKeown, K · 2019
Earlier work this paper cites.
Generative models from the perspective of continual learning
Lesort, T., Caselles-Dupré, H., Garcia-Ortiz, M., Goudou, J.-F., and Filliat, D · 2019
Earlier work this paper cites.
Bert post-training for review reading comprehension and aspect-based sentiment analysis
Xu, H., Liu, B., Shu, L., and Yu, P. S · 2019
Earlier work this paper cites.
Lifelong gan: Continual learning for conditional image generation
Zhai, M., Chen, L., Tung, F., He, J., Nawhal, M., and Mori, G · 2019
Earlier work this paper cites.
On warm-starting neural network training
Ash, J. and Adams, R. P · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Earlier work this paper cites.
Understanding the role of training regimes in continual learning
Mirzadeh, S. I., Farajtabar, M., Pascanu, R., and Ghasemzadeh, H · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Rasley, J., Rajbhandari, S., Ruwase, O., and He, Y · 2020
Cited alongside, same era.
Transformers: State-of-the-Art Natural Language Processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Le Scao, T., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M · 2020
Cited alongside, same era.
GPT-NeoX: Large Scale Autoregressive Language Modeling in PyTorch, 8 2021
Andonian, A., Anthony, Q., Biderman, S., Black, S., Gali, P., Gao, L., Hallahan, E., Levy-Kramer, J., Leahy, C., Nestler, L., Parker, K., Pieler, M., Purohit, S., Songz, T., Phil, W., and Weinbach, S · 2021
Cited alongside, same era.
On anytime learning at macroscale
Caccia, L., Xu, J., Ott, M., Ranzato, M., and Denoyer, L · 2021
Cited alongside, same era.
A framework for few-shot language model evaluation, September 2021
Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., Phang, J., Reynolds, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2021
Holistic evaluation of language models
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., et al · 2022
Later among the works it cites.
TimeLMs: Diachronic language models from Twitter
Loureiro, D., Barbieri, F., Neves, L., Espinosa Anke, L., and Camacho-collados, J · 2022
Later among the works it cites.
Continual learning with foundation models: An empirical study of latent replay, 2022
Ostapenko, O., Lesort, T., Rodríguez, P., Arefin, M. R., Douillard, A., Rish, I., and Charlin, L · 2022
Later among the works it cites.
Elle: Efficient lifelong pre-training for emerging data
Qin, Y., Zhang, J., Lin, Y., Liu, Z., Li, P., Sun, M., and Zhou, J · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
Scao, T. L., Fan, A., Akiki, C., Pavlick, E., Ilić, S., Hesslow, D., Castagné, R., Luccioni, A. S., Yvon, F., Gallé, M., et al · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Demix layers: Disentangling domains for modular language modeling
Gururangan, S., Lewis, M., Holtzman, A., Smith, N. A., and Zettlemoyer, L · 2021
Cited alongside, same era.
ECONET: Effective continual pretraining of language models for event temporal reasoning
Han, R., Ren, X., and Peng, N · 2021
Cited alongside, same era.
Towards continual knowledge learning of language models
Jang, J., Ye, S., Yang, S., Shin, J., Han, J., Kim, G., Choi, S. J., and Seo, M · 2021
Cited alongside, same era.
Understanding continual learning settings with data distribution drift analysis
Lesort, T., Caccia, M., and Rish, I · 2021
Cited alongside, same era.
Representational continuity for unsupervised continual learning
Madaan, D., Yoon, J., Li, Y., Liu, Y., and Hwang, S. J · 2021
Cited alongside, same era.
Scaling language models: Methods, analysis & insights from training gopher
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., et al · 2021
Cited alongside, same era.
Gpt-neox-20b: An open-source autoregressive language model, 2022
Black, S., Biderman, S., Hallahan, E., Anthony, Q., Gao, L., Golding, L., He, H., Leahy, C., McDonell, K., Phang, J., Pieler, M., Prashanth, U. S., Purohit, S., Reynolds, L., Tow, J., Wang, B., and Weinbach, S · 2022
Cited alongside, same era.
Later among the works it cites.
Fine-tuned language models are continual learners
Scialom, T., Chakrabarty, T., and Muresan, S · 2022
Later among the works it cites.
Tensor programs v: Tuning large neural networks via zero-shot hyperparameter transfer
Yang, G., Hu, E. J., Babuschkin, I., Sidor, S., Farhi, D., Pachocki, J., Liu, X., Chen, W., and Gao, J · 2022
Later among the works it cites.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S., Schoelkopf, H., Anthony, Q., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al · 2023
Closest in time.
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., Dollár, P., and Girshick, R · 2023
Closest in time.
Continual evaluation for lifelong learning: Identifying the stability gap
Lange, M. D., van de Ven, G. M., and Tuytelaars, T · 2023
Closest in time.
Challenging common assumptions about catastrophic forgetting
Lesort, T., Ostapenko, O., Rodriguez, P., Arefin, M. R., Misra, D., Charlin, L., and Rish, I · 2023
Closest in time.
Dinov2: Learning robust visual features without supervision, 2023
Oquab, M., Darcet, T., Moutakanni, T., Vo, H. V., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., Howes, R., Huang, P.-Y., Xu, H., Sharma, V., Li, S.-W., Galuba, W., Rabbat, M., Assran, M., Ballas, N., Synnaeve, G., Misra, I., Jegou, H., Mairal, J., Labatut, P., Joulin, A., and Bojanowski, P · 2023
Closest in time.
SlimPajama: A 627B token cleaned and deduplicated version of RedPajama
Soboleva, D., Al-Khateeb, F., Myers, R., Steeves, J. R., Hestness, J., and Dey, N · 2023
Closest in time.
Redpajama: An open source recipe to reproduce llama training dataset, 2023
Together.xyz · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Closest in time.
Overcoming catastrophic forgetting in massively multilingual continual learning
Winata, G. I., Xie, L., Radhakrishnan, K., Wu, S., Jin, X., Cheng, P., Kulkarni, M., and Preotiuc-Pietro, D · 2023
Closest in time.
A survey of large language models
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., et al · 2023
Closest in time.
NVIDIA Collective Communication Library (NCCL)
NVIDIA · 2026
Closest in time.
Pytorch extension with NVIDIA-maintained utilities to streamline mixed precision and distributed training
NVIDIA · 2026
Closest in time.