Fetching the paper…
Reading the bibliography…
Catastrophic forgetting emerges as a critical challenge when fine-tuning multi-modal large language models (MLLMs), where improving performance on unseen tasks often leads to a significant performance drop on the original tasks.
Optimal brain damage
LeCun, Y., Denker, J., and Solla, S · 1989
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
McCloskey, M. and Cohen, N. J · 1989
Earlier work this paper cites.
Optimal brain surgeon and general network pruning
Hassibi, B., Stork, D. G., and Wolff, G. J · 1993
Earlier work this paper cites.
An empirical investigation of catastrophic forgetting in gradient-based neural networks
Goodfellow, I. J., Mirza, M., Xiao, D., Courville, A., and Bengio, Y · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P., Lai, A., Hodosh, M., and Hockenmaier, J · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D · 2015
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
Memory aware synapses: Learning what (not) to forget
Aljundi, R., Babiloni, F., Elhoseiny, M., Rohrbach, M., and Tuytelaars, T · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
Gurari, D., Li, Q., Stangl, A. J., Guo, A., Lin, C., Grauman, K., Luo, J., and Bigham, J. P · 2018
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Howard, J. and Ruder, S · 2018
Earlier work this paper cites.
Online structured laplace approximations for overcoming catastrophic forgetting
Ritter, H., Botev, A., and Barber, D · 2018
Earlier work this paper cites.
Progress & compress: A scalable framework for continual learning
Schwarz, J., Czarnecki, W., Luketina, J., Grabska-Barwinska, A., Teh, Y. W., Pascanu, R., and Hadsell, R · 2018
Earlier work this paper cites.
Explicit inductive bias for transfer learning with convolutional networks
Xuhong, L., Grandvalet, Y., and Davoine, F · 2018
Earlier work this paper cites.
Nocaps: Novel object captioning at scale
Agrawal, H., Desai, K., Wang, Y., Chen, X., Jain, R., Johnson, M., Batra, D., Parikh, D., Lee, S., and Anderson, P · 2019
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A. and Manning, C. D · 2019
Earlier work this paper cites.
Mixout: Effective regularization to finetune large-scale pretrained language models
Lee, C., Cho, K., and Kang, W · 2019
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Marino, K., Rastegari, M., Farhadi, A., and Mottaghi, R · 2019
Earlier work this paper cites.
Towards vqa models that can read
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Recall and learn: Fine-tuning deep pretrained language models with less forgetting
Chen, S., Hou, Y., Cui, Y., Che, W., Liu, T., and Yu, X · 2020
Cited alongside, same era.
On the stability of fine-tuning bert: Misconceptions, explanations, and strong baselines
Mosbach, M., Andriushchenko, M., and Klakow, D · 2020
Cited alongside, same era.
Up or down? adaptive rounding for post-training quantization
Nagel, M., Amjad, R. A., Van Baalen, M., Louizos, C., and Blankevoort, T · 2020
Cited alongside, same era.
Towards language models that can see: Computer vision through the lens of natural language
Berrios, W., Mittal, G., Thrush, T., Kiela, D., and Singh, A · 2023
Later among the works it cites.
Leveraging large language models for scalable vector graphics-driven image understanding
Cai, M., Huang, Z., Li, Y., Wang, H., and Lee, Y. J · 2023
Later among the works it cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning. arxiv 2023
Dai, W., Li, J., Li, D., Tiong, A., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S · 2023
Later among the works it cites.
Sparsegpt: Massive language models can be accurately pruned in one-shot
Frantar, E. and Alistarh, D · 2023
Later among the works it cites.
Mme: A comprehensive evaluation benchmark for multimodal large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Few-shot class-incremental learning
Tao, X., Hong, X., Chang, X., Dong, S., Wei, X., and Gong, Y · 2020
Cited alongside, same era.
Towards accurate post-training network quantization via bit-split and stitching
Wang, P., Chen, Q., He, X., and Cheng, J · 2020
Cited alongside, same era.
Revisiting few-sample bert fine-tuning
Zhang, T., Wu, F., Katiyar, A., Weinberger, K. Q., and Artzi, Y · 2020
Cited alongside, same era.
How should pre-trained language models be fine-tuned towards adversarial robustness?
Dong, X., Luu, A. T., Lin, M., Yan, S., and Zhang, H · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
Accurate post training quantization with small calibration sets
Hubara, I., Nahshan, Y., Hanani, Y., Banner, R., and Soudry, D · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., Yang, J., Zheng, X., Li, K., Sun, X., et al · 2023
Later among the works it cites.
Continual instruction tuning for large multimodal models
He, J., Guo, H., Tang, M., and Wang, J · 2023
Later among the works it cites.
3d-llm: Injecting the 3d world into large language models
Hong, Y., Zhen, H., Chen, P., Zheng, S., Du, Y., Chen, Z., and Gan, C · 2023
Later among the works it cites.
Language is not all you need: Aligning perception with language models
Huang, S., Dong, L., Wang, W., Hao, Y., Singhal, S., Ma, S., Lv, T., Cui, L., Mohammed, O. K., Liu, Q., et al · 2023
Later among the works it cites.
Lin, Y., Tan, L., Lin, H., Zheng, Z., Pi, R., Zhang, J., Diao, S., Wang, H., Zhao, H., Yao, Y., et al · 2023
Later among the works it cites.
Task-specific skill localization in fine-tuned language models
Panigrahi, A., Saunshi, N., Zhao, H., and Arora, S · 2023
Later among the works it cites.
Ecoflap: Efficient coarse-to-fine layer-wise pruning for vision-language models
Sung, Y.-L., Yoon, J., and Bansal, M · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Voyager: An open-ended embodied agent with large language models
Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A · 2023
Later among the works it cites.
A brief overview of chatgpt: The history, status quo and potential future development
Wu, T., He, S., Liu, J., Sun, S., Liu, K., Han, Q.-L., and Tang, Y · 2023
Later among the works it cites.
Neural collapse inspired feature-classifier alignment for few-shot class incremental learning
Yang, Y., Yuan, H., Li, X., Lin, Z., Torr, P., and Tao, D · 2023
Later among the works it cites.
A survey on multimodal large language models
Yin, S., Fu, C., Zhao, S., Li, K., Sun, X., Xu, T., and Chen, E · 2023
Later among the works it cites.
Language models are super mario: Absorbing abilities from homologous models as a free lunch
Yu, L., Yu, B., Yu, H., Huang, F., and Li, Y · 2023
Later among the works it cites.
Investigating the catastrophic forgetting in multimodal large language models
Zhai, Y., Tong, S., Li, X., Cai, M., Qu, Q., Lee, Y. J., and Ma, Y · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2023
Later among the works it cites.