Fetching the paper…
Reading the bibliography…
The recent increase in data and model scale for language model pre-training has led to huge training costs.
French, R.M.: Catastrophic forgetting in connectionist networks. Trends in Cognitive Sciences pp. 128–135 (1999)
1999
Earlier work this paper cites.
2016
Earlier work this paper cites.
Littell, P., Mortensen, D.R., Lin, K., Kairis, K., Turner, C., Levin, L.: URIEL and lang2vec: Representing languages as typological, geographical, and phylogenetic vectors. In: Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers. pp. 8–14 (2017)
2017
Earlier work this paper cites.
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al.: Improving language understanding by generative pre-training. OpenAI (2018)
2018
Earlier work this paper cites.
Aljundi, R., Belilovsky, E., Tuytelaars, T., Charlin, L., Caccia, M., Lin, M., Page-Caccia, L.: Online continual learning with maximal interfered retrieval. In: Advances in Neural Information Processing Systems (2019)
2019
Earlier work this paper cites.
de Masson d'Autume, C., Ruder, S., Kong, L., Yogatama, D.: Episodic memory in lifelong language learning. In: Advances in Neural Information Processing Systems (2019)
2019
Earlier work this paper cites.
Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: BERT: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). pp. 4171–4186 (Jun 2019)
2019
Earlier work this paper cites.
Parisi, G.I., Kemker, R., Part, J.L., Kanan, C., Wermter, S.: Continual lifelong learning with neural networks: A review. Neural Networks pp. 54 – 71 (2019)
2019
Earlier work this paper cites.
Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., Grave, É., Ott, M., Zettlemoyer, L., Stoyanov, V.: Unsupervised cross-lingual representation learning at scale. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. pp. 8440–8451 (2020)
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Magdalena Biesialska, Katarzyna Biesialska, M.R.C.j.: Continual lifelong learning in natural language processing: A survey. In: Proceedings of the 28th International Conference on Computational Linguistics. pp. 6523–6541 (2020)
2020
Earlier work this paper cites.
Ortiz Suárez, P.J., Romary, L., Sagot, B.: A monolingual approach to contextualized word embeddings for mid-resource languages. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. pp. 1703–1714 (2020)
2020
Earlier work this paper cites.
Abnar, S., Dehghani, M., Neyshabur, B., Sedghi, H.: Exploring the limits of large scale pre-training. In: International Conference on Learning Representations (2021)
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Cited alongside, same era.
Han, R., Ren, X., Peng, N.: ECONET: Effective continual pretraining of language models for event temporal reasoning. In: EMNLP 2021. pp. 5367–5380 (2021)
2021
Cited alongside, same era.
Lazaridou, A., Kuncoro, A., Gribovskaya, E., Agrawal, D., Liska, A., Terzi, T., Gimenez, M., de Masson d'Autume, C., Kocisky, T., Ruder, S., Yogatama, D., Cao, K., Young, S., Blunsom, P.: Mind the gap: Assessing temporal generalization in neural language models. In: Advances in Neural Information Processing Systems. pp. 29348–29363 (2021)
2021
Cited alongside, same era.
2021
Loureiro, D., Barbieri, F., Neves, L., Espinosa Anke, L., Camacho-collados, J.: TimeLMs: Diachronic language models from Twitter. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations (2022)
2022
Later among the works it cites.
Mao, Y., Liang, Y., Duan, N., Wang, H., Wang, K., Chen, L., Gao, Y.: Less-forgetting multi-lingual fine-tuning. In: Advances in Neural Information Processing Systems. pp. 14917–14928 (2022)
2022
Later among the works it cites.
Ramasesh, V.V., Lewkowycz, A., Dyer, E.: Effect of scale on catastrophic forgetting in neural networks. In: International Conference on Learning Representations (2022)
2022
Later among the works it cites.
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Strømberg-Derczynski, L., Ciosici, M., Baglini, R., Christiansen, M.H., Dalsgaard, J.A., Fusaroli, R., Henrichsen, P.J., Hvingelby, R., Kirkedal, A., Kjeldsen, A.S., Ladefoged, C., Nielsen, F.Å., Madsen, J., Petersen, M.L., Rystrøm, J.H., Varab, D.: The Danish Gigaword corpus. In: Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa). pp. 413–421 (2021)
2021
Cited alongside, same era.
Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A., Raffel, C.: mT5: A massively multilingual pre-trained text-to-text transformer. In: Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. pp. 483–498 (2021)
2021
Cited alongside, same era.
2022
Cited alongside, same era.
Barkarson, S., Steingrímsson, S., Hafsteinsdóttir, H.: Evolving large text corpora: Four versions of the Icelandic Gigaword Corpus. In: Proceedings of the Thirteenth Language Resources and Evaluation Conference. pp. 2371–2381 (2022)
2022
Cited alongside, same era.
Blevins, T., Zettlemoyer, L.: Language contamination helps explains the cross-lingual capabilities of English pretrained models. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (2022)
2022
Cited alongside, same era.
Coria, J.M., Veron, M., Ghannay, S., Bernard, G., Bredin, H., Galibert, O., Rosset, S.: Analyzing BERT cross-lingual transfer capabilities in continual sequence labeling. In: Proceedings of the First Workshop on Performance and Interpretability Evaluations of Multimodal, Multipurpose, Massive-Scale Models. pp. 15–25. International Conference on Computational Linguistics (2022)
2022
Cited alongside, same era.
Jang, J., Ye, S., Lee, C., Yang, S., Shin, J., Han, J., Kim, G., Seo, M.: TemporalWiki: A lifelong benchmark for training and evaluating ever-evolving language models. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. pp. 6237–6250 (Dec 2022)
2022
Cited alongside, same era.
Kummervold, P., Wetjen, F., de la Rosa, J.: The Norwegian colossal corpus: A text corpus for training large Norwegian language models. In: Proceedings of the Thirteenth Language Resources and Evaluation Conference. pp. 3852–3860 (Jun 2022)
2022
Cited alongside, same era.
Scialom, T., Chakrabarty, T., Muresan, S.: Fine-tuned language models are continual learners. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. pp. 6107–6122 (2022)
2022
Later among the works it cites.
2023
Closest in time.
Ke, Z., Shao, Y., Lin, H., Konishi, T., Kim, G., Liu, B.: Continual pre-training of language models. In: The Eleventh International Conference on Learning Representations (2023)
2023
Closest in time.
Lesort, T., Ostapenko, O., Rodríguez, P., Misra, D., Arefin, M.R., Charlin, L., Rish, I.: Challenging common assumptions about catastrophic forgetting and knowledge accumulation. In: Chandar, S., Pascanu, R., Sedghi, H., Precup, D. (eds.) Proceedings of The 2nd Conference on Lifelong Learning Agents. Proceedings of Machine Learning Research, vol. 232, pp. 43–65. PMLR (22–25 Aug 2023), https://proceedings.mlr.press/v232/lesort23a.html
2023
Closest in time.
2023
Closest in time.
Philippy, F., Guo, S., Haddadan, S.: Towards a common understanding of contributing factors for cross-lingual transfer in multilingual language models: A review. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). pp. 5877–5891 (2023)
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.