Learn to grow: A continual structure learning framework for overcoming catastrophic forgetting
Li, X., Zhou, Y., Wu, T., Socher, R., and Xiong, C. (2019) · 2019
Later among the works it cites.
Continual lifelong learning with neural networks: A review
Parisi, G. I., Kemker, R., Part, J. L., Kanan, C., and Wermter, S. (2019) · 2019
Later among the works it cites.
Experience replay for continual learning
Rolnick, D., Ahuja, A., Schwarz, J., Lillicrap, T. P., and Wayne, G. (2019) · 2019
Later among the works it cites.
An empirical study of example forgetting during deep neural network learning
Toneva, M., Sordoni, A., Combes, R. T. d., Trischler, A., Bengio, Y., and Gordon, G. J. (2019) · 2019
Later among the works it cites.
Learning to continually learn
Beaulieu, S., Frati, L., Miconi, T., Lehman, J., Stanley, K. O., Clune, J., and Cheney, N. (2020) · 2020
Later among the works it cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 2020
Later among the works it cites.
Continual learning in low-rank orthogonal subspaces
Chaudhry, A., Khan, N., Dokania, P. K., and Torr, P. H. (2020) · 2020
Later among the works it cites.
Orthogonal gradient descent for continual learning
Farajtabar, M., Azizan, N., Mott, A., and Li, A. (2020) · 2020
Later among the works it cites.
Understanding the role of training regimes in continual learning
Mirzadeh, S. I., Farajtabar, M., Pascanu, R., and Ghasemzadeh, H. (2020b) · 2020
Later among the works it cites.
Supermasks in superposition
Wortsman, M., Ramanujan, V., Liu, R., Kembhavi, A., Rastegari, M., Yosinski, J., and Farhadi, A. (2020) · 2020
Later among the works it cites.
Gradient descent optimizes over-parameterized deep relu networks
Zou, D., Cao, Y., Zhou, D., and Gu, Q. (2020) · 2020
Later among the works it cites.
Using hindsight to anchor past knowledge in continual learning
Chaudhry, A., Gordo, A., Dokania, P. K., Torr, P. H. S., and Lopez-Paz, D. (2021) · 2021
Closest in time.
A theoretical analysis of catastrophic forgetting through the ntk overlap matrix
Doan, T., Bennani, M. A., Mazoure, B., Rabusseau, G., and Alquier, P. (2021) · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. (2021) · 2021
Closest in time.
Dynamic language models for continuously evolving content
Hombaiah, S. A., Chen, T., Zhang, M., Bendersky, M., and Najork, M. (2021) · 2021
Closest in time.
Task-agnostic continual learning with hybrid probabilistic models
Kirichenko, P., Farajtabar, M., Rao, D., Lakshminarayanan, B., Levine, N., Li, A., Hu, H., Wilson, A. G., and Pascanu, R. (2021) · 2021
Closest in time.
Pitfalls of static language modelling
Original
Lazaridou, A., Kuncoro, A., Gribovskaya, E., Agrawal, D., Liska, A., Terzi, T., Gimenez, M., d’Autume, C. d. M., Ruder, S., Yogatama, D., et al. (2021) · 2021
Closest in time.
An empirical investigation of the role of pre-training in lifelong learning
Mehta, S. V., Patil, D., Chandar, S., and Strubell, E. (2021) · 2021
Closest in time.
Linear mode connectivity in multitask and continual learning
Mirzadeh, S. I., Farajtabar, M., Gorur, D., Pascanu, R., and Ghasemzadeh, H. (2021) · 2021
Closest in time.
Efficient continual learning with modular networks and task-driven priors
Veniat, T., Denoyer, L., and Ranzato, M. (2021) · 2021
Closest in time.
Architecture matters in continual learning
Original
Mirzadeh, S. I., Chaudhry, A., Yin, D., Nguyen, T., Pascanu, R., Gorur, D., and Farajtabar, M. (2022) · 2022
Closest in time.
Effect of scale on catastrophic forgetting in neural networks
Ramasesh, V. V., Lewkowycz, A., and Dyer, E. (2022) · 2022
Closest in time.