Fetching the paper…
Reading the bibliography…
Backward compatibility of model predictions is a desired property when updating a machine learning driven application.
Catastrophic interference in connectionist networks: The sequential learning problem
McCloskey, M. and Cohen, N. J · 1989
Earlier work this paper cites.
Connectionist models of recognition memory: Constraints imposed by learning and forgetting functions
Ratcliff, R · 1990
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G. E., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Zhang, X., Zhao, J., and LeCun, Y · 2015
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N. C., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., Hassabis, D., Clopath, C., Kumaran, D., and Hadsell, R · 2017
Earlier work this paper cites.
Gradient episodic memory for continual learning
Lopez-Paz, D. and Ranzato, M · 2017
Earlier work this paper cites.
icarl: Incremental classifier and representation learning
Rebuffi, S.-A., Kolesnikov, A., Sperl, G., and Lampert, C. H · 2017
Earlier work this paper cites.
Neural domain adaptation for biomedical question answering
Wiese, G., Weissenborn, D., and Neves, M · 2017
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
Izmailov, P., Wilson, A., Podoprikhin, D., Vetrov, D., and Garipov, T · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Earlier work this paper cites.
Continual lifelong learning with neural networks: A review
Parisi, G. I., Kemker, R., Part, J. L., Kanan, C., and Wermter, S · 2019
Cited alongside, same era.
Efficient intent detection with dual sentence encoders
Casanueva, I., Temčinas, T., Gerz, D., Henderson, M., and Vulić, I · 2020
Cited alongside, same era.
Improved schemes for episodic memory-based lifelong learning
Guo, Y., Liu, M., Yang, T., and Rosing, T · 2020
Cited alongside, same era.
Mixout: Effective regularization to finetune large-scale pretrained language models
Lee, C., Cho, K., and Kang, W · 2020
Cited alongside, same era.
Towards backward-compatible representation learning
Shen, Y., Xiong, Y., Xia, W., and Soatto, S · 2020
Cited alongside, same era.
Editing factual knowledge in language models
De Cao, N., Aziz, W., and Titov, I · 2021
Git re-basin: Merging models modulo permutation symmetries
Ainsworth, S. K., Hayase, J., and Srinivasa, S. S · 2022
Later among the works it cites.
BitFit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Ben Zaken, E., Goldberg, Y., and Ravfogel, S · 2022
Later among the works it cites.
Measuring and reducing model update regression in structured prediction for nlp
Cai, D., Mansimov, E., Lai, Y.-A., Su, Y., Shu, L., and Zhang, Y · 2022
Later among the works it cites.
FitzGerald, J. G. M., Hench, C. L., Peris, C. S., Mackie, S., Rottmann, K., Sánchez, A. P. D., Nash, A., Urbach, L., Kakarala, V., Singh, R., Ranganath, S., Crist, L., Britan, M., Leeuwis, W., Tur, G., and Natarajan, P · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Merging models with fisher-weighted averaging
Matena, M. and Raffel, C · 2021
Cited alongside, same era.
On the stability of fine-tuning {bert}: Misconceptions, explanations, and strong baselines
Mosbach, M., Andriushchenko, M., and Klakow, D · 2021
Cited alongside, same era.
Regression bugs are in your model! measuring, reducing and analyzing regressions in NLP model updates
Xie, Y., Lai, Y.-A., Xiong, Y., Zhang, Y., and Soatto, S · 2021
Cited alongside, same era.
Positive-congruent training: Towards regression-free model updates
Yan, S., Xiong, Y., Kundu, K., Yang, S., Deng, S., Wang, M., Xia, W., and Soatto, S · 2021
Cited alongside, same era.
Elodi: Ensemble logit difference inhibition for positive-congruent training
Zhao, Y., Shen, Y., Xiong, Y., Yang, S., Xia, W., Tu, Z., Shiele, B., and Soatto, S. · 2021
Cited alongside, same era.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Wortsman, M., Ilharco, G., Gadre, S. Y., Roelofs, R., Lopes, R. G., Morcos, A. S., Namkoong, H., Farhadi, A., Carmon, Y., Kornblith, S., and Schmidt, L
Cited in the paper.
Ilharco, G., Wortsman, M., Gadre, S. Y., Song, S., Hajishirzi, H., Kornblith, S., Farhadi, A., and Schmidt, L · 2022
Later among the works it cites.
A continual learning survey: Defying forgetting in classification tasks
Lange, M. D., Aljundi, R., Masana, M., Parisot, S., Jia, X., Leonardis, A., Slabaugh, G. G., and Tuytelaars, T · 2022
Later among the works it cites.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning
Liu, H., Tam, D., Mohammed, M., Mohta, J., Huang, T., Bansal, M., and Raffel, C · 2022
Later among the works it cites.
Fast model editing at scale
Mitchell, E., Lin, C., Bosselut, A., Finn, C., and Manning, C. D · 2022
Later among the works it cites.
Diverse weight averaging for out-of-distribution generalization
Rame, A., Kirchmeyer, M., Rahier, T., Rakotomamonjy, A., patrick gallinari, and Cord, M · 2022
Later among the works it cites.
An introduction to lifelong supervised learning
Sodhani, S., Farmazi, M., Mehta, S. V., Malviya, P., Abdelsalam, M., Janarthanan, J., and Chandar, S · 2022
Later among the works it cites.
Robust fine-tuning of zero-shot models
Wortsman, M., Ilharco, G., Li, M., Kim, J. W., Hajishirzi, H., Farhadi, A., Namkoong, H., and Schmidt, L · 2022
Later among the works it cites.