Fetching the paper…
Reading the bibliography…
Multimodal Continual Instruction Tuning (MCIT) enables Multimodal Large Language Models (MLLMs) to meet continuously emerging requirements without expensive retraining.
G. H. Golub and C. Reinsch, “Singular value decomposition and least squares solutions,” in Handbook for Automatic Computation: Volume II: Linear Algebra . Springer, 1971, pp. 134–151
1971
Earlier work this paper cites.
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier, “From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,” Transactions of the Association for Computational Linguistics , vol. 2, pp. 67–78, 2014
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 3128–3137
2015
Earlier work this paper cites.
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the v in vqa matter: Elevating the role of image understanding in visual question answering,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 6904–6913
2017
Earlier work this paper cites.
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska et al. , “Overcoming catastrophic forgetting in neural networks,” Proceedings of the national academy of sciences , vol. 114, no. 13, pp. 3521–3526, 2017
2017
Earlier work this paper cites.
D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham, “Vizwiz grand challenge: Answering visual questions from blind people,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 3608–3617
2018
Earlier work this paper cites.
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach, “Towards vqa models that can read,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 8317–8326
2019
Earlier work this paper cites.
D. A. Hudson and C. D. Manning, “Gqa: A new dataset for real-world visual reasoning and compositional question answering,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 6700–6709
2019
Earlier work this paper cites.
A. Mishra, S. Shekhar, A. K. Singh, and A. Chakraborty, “Ocr-vqa: Visual question answering by reading text in images,” in 2019 international conference on document analysis and recognition (ICDAR) . IEEE, 2019, pp. 947–952
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi, “Ok-vqa: A visual question answering benchmark requiring external knowledge,” in Proceedings of the IEEE/cvf conference on computer vision and pattern recognition , 2019, pp. 3195–3204
2019
Earlier work this paper cites.
A. Chaudhry, M. Ranzato, M. Rohrbach, and M. Elhoseiny, “Efficient lifelong learning with a-gem,” in International Conference on Learning Representations , 2019
2019
Earlier work this paper cites.
O. Sidorov, R. Hu, M. Rohrbach, and A. Singh, “Textcaps: a dataset for image captioning with reading comprehension,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16 . Springer, 2020, pp. 742–758
2020
Earlier work this paper cites.
P. Buzzega, M. Boschini, A. Porrello, D. Abati, and S. Calderara, “Dark experience for general continual learning: a strong, simple baseline,” Advances in neural information processing systems , vol. 33, pp. 15 920–15 930, 2020
2020
Earlier work this paper cites.
G. Saha, I. Garg, and K. Roy, “Gradient projection memory for continual learning,” in International Conference on Learning Representations , 2021
2021
Cited alongside, same era.
Z. Wang, Z. Zhang, C.-Y. Lee, H. Zhang, R. Sun, X. Ren, G. Su, V. Perot, J. Dy, and T. Pfister, “Learning to prompt for continual learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 139–149
2022
Cited alongside, same era.
Z. Wang, Z. Zhang, S. Ebrahimi, R. Sun, H. Zhang, C.-Y. Lee, X. Ren, G. Su, V. Perot, J. Dy et al. , “Dualprompt: Complementary prompting for rehearsal-free continual learning,” in European Conference on Computer Vision . Springer, 2022, pp. 631–648
2022
Cited alongside, same era.
T. Srinivasan, T.-Y. Chang, L. Pinto Alva, G. Chochlakis, M. Rostami, and J. Thomason, “Climb: A continual learning benchmark for vision-and-language tasks,” Advances in Neural Information Processing Systems , vol. 35, pp. 29 440–29 453, 2022
2022
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
J. S. Smith, L. Karlinsky, V. Gutta, P. Cascante-Bonilla, D. Kim, A. Arbelle, R. Panda, R. Feris, and Z. Kira, “Coda-prompt: Continual decomposed attention-based prompting for rehearsal-free continual learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 11 909–11 919
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
D. Schwenk, A. Khandelwal, C. Clark, K. Marino, and R. Mottaghi, “A-okvqa: A benchmark for visual question answering using world knowledge,” in European Conference on Computer Vision . Springer, 2022, pp. 146–162
2022
Cited alongside, same era.
D. Lopez-Paz and M. Ranzato, “Gradient episodic memory for continual learning,” Advances in neural information processing systems , vol. 30, 2017
2022
Cited alongside, same era.
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi, “Instructblip: Towards general-purpose vision-language models with instruction tuning,” 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
OpenAI, “Gpt-4 technical report,” 2023
2023
Cited alongside, same era.
2023
Later among the works it cites.
J. Qiao, X. Tan, C. Chen, Y. Qu, Y. Peng, Y. Xie et al. , “Prompt gradient projection for continual learning,” in The Twelfth International Conference on Learning Representations , 2023
2023
Later among the works it cites.
X. Zhang, F. Zhang, and C. Xu, “Vqacl: A novel visual question answering continual learning setting,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 19 102–19 112
2023
Later among the works it cites.
Z. Qian, X. Wang, X. Duan, P. Qin, Y. Li, and W. Zhu, “Decouple before interact: Multi-modal prompt learning for continual visual question answering,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 2953–2962
2023
Later among the works it cites.
Z. Ni, L. Wei, S. Tang, Y. Zhuang, and Q. Tian, “Continual vision-language representation learning with off-diagonal information,” 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
H. Zhu, Y. Wei, X. Liang, C. Zhang, and Y. Zhao, “Ctp: Towards vision-language continual pretraining via compatible momentum contrast and topology preservation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 22 257–22 267
2023
Later among the works it cites.
Y. Fang, W. Wang, B. Xie, Q. Sun, L. Wu, X. Wang, T. Huang, X. Wang, and Y. Cao, “Eva: Exploring the limits of masked visual representation learning at scale,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 19 358–19 369
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
D. Li, J. Li, H. Le, G. Wang, S. Savarese, and S. C. Hoi, “Lavis: A one-stop library for language-vision intelligence,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations) , 2023, pp. 31–41
2023
Later among the works it cites.