Fetching the paper…
Reading the bibliography…
Continual learning (CL) has garnered significant attention because of its ability to adapt to new tasks that arrive over time.
Y. LeCun, B. Boser, J. Denker, D. Henderson, R. Howard, W. Hubbard, and L. Jackel, “Handwritten digit recognition with a back-propagation network,”
1989
Earlier work this paper cites.
M. McCloskey and N. J. Cohen, “Catastrophic interference in connectionist networks: The sequential learning problem,” in
1989
Earlier work this paper cites.
K. Viele and B. Tong, “Modeling with mixtures of linear regressions,”
2002
Earlier work this paper cites.
A. Krizhevsky, G. Hinton
2009
Earlier work this paper cites.
2013
Earlier work this paper cites.
A. Anandkumar, R. Ge, D. J. Hsu, S. M. Kakade, M. Telgarsky
2014
Earlier work this paper cites.
Y. Le and X. Yang, “Tiny imagenet visual recognition challenge,”
2015
Earlier work this paper cites.
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, “Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,” in
2016
Earlier work this paper cites.
K. Zhong, P. Jain, and I. S. Dhillon, “Mixed linear regression with multiple components,”
2016
Earlier work this paper cites.
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska
2017
Earlier work this paper cites.
M. Belkin, S. Ma, and S. Mandal, “To understand deep learning we need to understand kernel learning,” in
2018
Earlier work this paper cites.
A. Chaudhry, M. Ranzato, M. Rohrbach, and M. Elhoseiny, “Efficient lifelong learning with a-gem,”
2018
Earlier work this paper cites.
S. Gunasekar, J. Lee, D. Soudry, and N. Srebro, “Characterizing implicit bias in terms of optimization geometry,” in
2018
Earlier work this paper cites.
H. D. Nguyen and F. Chamroukhi, “Practical and theoretical aspects of mixture-of-experts modeling: An overview,”
2018
Earlier work this paper cites.
H. Ritter, A. Botev, and D. Barber, “Online structured laplace approximations for overcoming catastrophic forgetting,”
2018
Earlier work this paper cites.
J. Serra, D. Suris, M. Miron, and A. Karatzoglou, “Overcoming catastrophic forgetting with hard attention to the task,” in
2018
Earlier work this paper cites.
G. Jerfel, E. Grant, T. Griffiths, and K. A. Heller, “Reconciling meta-learning and continual learning with online mixtures of tasks,”
2019
Earlier work this paper cites.
G. I. Parisi, R. Kemker, J. L. Part, C. Kanan, and S. Wermter, “Continual lifelong learning with neural networks: A review,”
2019
Earlier work this paper cites.
J. Yoon, S. Kim, E. Yang, and S. J. Hwang, “Scalable and order-robust continual learning with additive parameter decomposition,” in
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
M. Farajtabar, N. Azizan, A. Mott, and A. Li, “Orthogonal gradient descent for continual learning,” in
2020
Cited alongside, same era.
P. Ju, X. Lin, and J. Liu, “Overfitting can be harmless for basis pursuit, but only to a degree,”
2020
Cited alongside, same era.
S. Lee, J. Ha, D. Zhang, and G. Kim, “A neural dirichlet process mixture model for task-free continual learning,” in
2020
Cited alongside, same era.
G. Saha, I. Garg, and K. Roy, “Gradient projection memory for continual learning,” in
2020
Cited alongside, same era.
T. Doan, M. A. Bennani, B. Mazoure, G. Rabusseau, and P. Alquier, “A theoretical analysis of catastrophic forgetting through the ntk overlap matrix,” in
2021
Cited alongside, same era.
W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,”
2022
Later among the works it cites.
L. Wang, X. Zhang, Q. Li, J. Zhu, and Y. Zhong, “Coscl: Cooperation of small continual learners is stronger than a big one,” in
2022
Later among the works it cites.
Q. Wang and H. Van Hoof, “Learning expressive meta-representations with mixture of expert neural processes,”
2022
Later among the works it cites.
Y. Zhou, T. Lei, H. Liu, N. Du, Y. Huang, V. Zhao, A. M. Dai, Q. V. Le, J. Laudon
2022
Later among the works it cites.
T. Doan, S. I. Mirzadeh, and M. Farajtabar, “Continual learning beyond a single model,” in
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
H. Hihn and D. A. Braun, “Mixture-of-variational-experts for continual learning,”
2021
Cited alongside, same era.
X. Jin, A. Sadhu, J. Du, and X. Ren, “Gradient-based editing of memory examples for online task-free continual learning,”
2021
Cited alongside, same era.
S. Lee, S. Goldt, and A. Saxe, “Continual learning in the teacher-student setup: Impact of task similarity,” in
2021
Cited alongside, same era.
S. Lin, L. Yang, D. Fan, and J. Zhang, “Trgp: Trust region gradient projection for continual learning,” in
2021
Cited alongside, same era.
H. Liu and H. Liu, “Continual learning with recursive gradient optimization,” in
2021
Cited alongside, same era.
C. Riquelme, J. Puigcerver, B. Mustafa, M. Neumann, R. Jenatton, A. Susano Pinto, D. Keysers, and N. Houlsby, “Scaling vision with sparse mixture of experts,”
2021
Cited alongside, same era.
R. Gao and W. Liu, “Ddgr: Continual learning with deep diffusion-based generative replay,” in
2023
Later among the works it cites.
T. Konishi, M. Kurokawa, C. Ono, Z. Ke, G. Kim, and B. Liu, “Parameter-level soft-masking for continual learning,” in
2023
Later among the works it cites.
T. Lesort, O. Ostapenko, P. Rodríguez, D. Misra, M. R. Arefin, L. Charlin, and I. Rish, “Challenging common assumptions about catastrophic forgetting and knowledge accumulation,” in
2023
Later among the works it cites.
S. Lin, P. Ju, Y. Liang, and N. Shroff, “Theory on forgetting and generalization of continual learning,” in
2023
Later among the works it cites.
L. Peng, P. Giampouras, and R. Vidal, “The ideal continual learner: An agent that never forgets,” in
2023
Later among the works it cites.
G. Rypeść, S. Cygert, V. Khan, T. Trzcinski, B. M. Zieliński, and B. Twardowski, “Divide and not forget: Ensemble of selectively trained experts in continual learning,” in
2023
Later among the works it cites.
T. Zadouri, A. Üstün, A. Ahmadian, B. Ermis, A. Locatelli, and S. Hooker, “Pushing mixture of experts to the limit: Extremely parameter efficient moe for instruction tuning,” in
2023
Later among the works it cites.
Y. Huang, Y. Cheng, and Y. Liang, “In-context convergence of transformers,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
H. Nguyen, P. Akbarian, F. Yan, and N. Ho, “Statistical perspective of top-k sparse softmax gating mixture of experts,” in
2024
Closest in time.
L. Wang, X. Zhang, H. Su, and J. Zhu, “A comprehensive survey of continual learning: Theory, method and application,”
2024
Closest in time.
2024
Closest in time.