Reinforced continual learning
Xu, J. and Zhu, Z. (2018) · 2018
Later among the works it cites.
Lifelong learning with dynamically expandable networks
Yoon, J., Yang, E., Lee, J., and Hwang, S. J. (2018) · 2018
Later among the works it cites.
Learn to grow: A continual structure learning framework for overcoming catastrophic forgetting
Li, X., Zhou, Y., Wu, T., Socher, R., and Xiong, C. (2019) · 2019
Later among the works it cites.
Experience replay for continual learning
Rolnick, D., Ahuja, A., Schwarz, J., Lillicrap, T. P., and Wayne, G. (2019) · 2019
Later among the works it cites.
Orthogonal gradient descent for continual learning
Farajtabar, M., Azizan, N., Mott, A., and Li, A. (2020) · 2020
Later among the works it cites.
Understanding the role of training regimes in continual learning
Mirzadeh, S. I., Farajtabar, M., Pascanu, R., and Ghasemzadeh, H. (2020) · 2020
Later among the works it cites.
Supermasks in superposition
Wortsman, M., Ramanujan, V., Liu, R., Kembhavi, A., Rastegari, M., Yosinski, J., and Farhadi, A. (2020) · 2020
Later among the works it cites.
Coatnet: Marrying convolution and attention for all data sizes
Dai, Z., Liu, H., Le, Q. V., and Tan, M. (2021) · 2021
Later among the works it cites.
Kernel continual learning
Derakhshani, M. M., Zhen, X., Shao, L., and Snoek, C. (2021) · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. (2021) · 2021
Later among the works it cites.
Task-agnostic continual learning with hybrid probabilistic models
Kirichenko, P., Farajtabar, M., Rao, D., Lakshminarayanan, B., Levine, N., Li, A., Hu, H., Wilson, A. G., and Pascanu, R. (2021) · 2021
Later among the works it cites.
Sharing less is more: Lifelong learning in deep networks with selective layer transfer
Lee, S., Behpour, S., and Eaton, E. (2021) · 2021
Later among the works it cites.
Vision transformers are robust learners
Paul, S. and Chen, P.-Y. (2022) · 2022
Closest in time.
Continual normalization: Rethinking batch normalization for online continual learning
Pham, Q., Liu, C., and HOI, S. (2022) · 2022
Closest in time.
Effect of scale on catastrophic forgetting in neural networks
Ramasesh, V. V., Lewkowycz, A., and Dyer, E. (2022) · 2022
Closest in time.