Fetching the paper…
Reading the bibliography…
Instruction tuning represents a prevalent strategy employed by Multimodal Large Language Models (MLLMs) to align with human instructions and adapt to new tasks.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. L. Berg · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Learning without forgetting
Z. Li and D. Hoiem · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy · 2016
Earlier work this paper cites.
Progressive neural networks
A. A. Rusu, N. C. Rabinowitz, G. Desjardins, H. Soyer, J. Kirkpatrick, K. Kavukcuoglu, R. Pascanu, and R. Hadsell · 2016
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
Gradient episodic memory for continual learning
D. Lopez-Paz and M. Ranzato · 2017
Earlier work this paper cites.
icarl: Incremental classifier and representation learning
S.-A. Rebuffi, A. Kolesnikov, G. Sperl, and C. H. Lampert · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. V. Le, G. E. Hinton, and J. Dean · 2017
Earlier work this paper cites.
Continual learning through synaptic intelligence
F. Zenke, B. Poole, and S. Ganguli · 2017
Earlier work this paper cites.
Memory aware synapses: Learning what (not) to forget
R. Aljundi, F. Babiloni, M. Elhoseiny, M. Rohrbach, and T. Tuytelaars · 2018
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham · 2018
Earlier work this paper cites.
Packnet: Adding multiple tasks to a single network by iterative pruning
A. Mallya and S. Lazebnik · 2018
Earlier work this paper cites.
Progress & compress: A scalable framework for continual learning
J. Schwarz, W. Czarnecki, J. Luketina, A. Grabska-Barwinska, Y. W. Teh, R. Pascanu, and R. Hadsell · 2018
Earlier work this paper cites.
Overcoming catastrophic forgetting with hard attention to the task
J. Serrà, D. Suris, M. Miron, and A. Karatzoglou · 2018
Earlier work this paper cites.
Efficient lifelong learning with A-GEM
A. Chaudhry, M. Ranzato, M. Rohrbach, and M. Elhoseiny · 2019
Earlier work this paper cites.
GQA: A new dataset for real-world visual reasoning and compositional question answering
D. A. Hudson and C. D. Manning · 2019
Earlier work this paper cites.
OCR-VQA: visual question answering by reading text in images
A. Mishra, S. Shekhar, A. K. Singh, and A. Chakraborty · 2019
Earlier work this paper cites.
Learning to learn without forgetting by maximizing transfer and minimizing interference
M. Riemer, I. Cases, R. Ajemian, M. Liu, I. Rish, Y. Tu, and G. Tesauro · 2019
Earlier work this paper cites.
Towards VQA models that can read
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach · 2019
Earlier work this paper cites.
Dark experience for general continual learning: a strong, simple baseline
P. Buzzega, M. Boschini, A. Porrello, D. Abati, and S. Calderara · 2020
Cited alongside, same era.
Supermasks in superposition
M. Wortsman, V. Ramanujan, R. Liu, A. Kembhavi, M. Rastegari, J. Yosinski, and A. Farhadi · 2020
Cited alongside, same era.
EEC: learning to encode and regenerate images for continual learning
A. Ayub and A. R. Wagner · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2021
Cited alongside, same era.
Instructblip: Towards general-purpose vision-language models with instruction tuning
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. C. H. Hoi · 2023
Later among the works it cites.
Loramoe: Revolutionizing mixture of experts for maintaining world knowledge in language model alignment
S. Dou, E. Zhou, Y. Liu, S. Gao, J. Zhao, W. Shen, Y. Zhou, Z. Xi, X. Wang, X. Fan, S. Pu, J. Zhu, R. Zheng, T. Gui, Q. Zhang, and X. Huang · 2023
Later among the works it cites.
Tinystories: How small can language models be and still speak coherent english?
R. Eldan and Y. Li · 2023
Later among the works it cites.
S. Gunasekar, Y. Zhang, J. Aneja, C. C. T. Mendes, A. D. Giorno, S. Gopi, M. Javaheripi, P. Kauffmann, G. de Rosa, O. Saarikivi, A. Salim, S. Shah, H. S. Behl, X. Wang, S. Bubeck, R. Eldan, A. T. Kalai, Y. T. Lee, and Y. Li · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Cited alongside, same era.
How many data points is a prompt worth?
T. L. Scao and A. M. Rush · 2021
Cited alongside, same era.
Efficient continual learning with modular networks and task-driven priors
T. Veniat, L. Denoyer, and M. Ranzato · 2021
Cited alongside, same era.
Finetuned language models are zero-shot learners
J. Wei, M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le · 2021
Cited alongside, same era.
New insights on reducing abrupt representation change in online continual learning
L. Caccia, R. Aljundi, N. Asadi, T. Tuytelaars, J. Pineau, and E. Belilovsky · 2022
Cited alongside, same era.
Class gradient projection for continual learning
C. Chen, J. Zhang, J. Song, and L. Gao · 2022
Cited alongside, same era.
H. Gupta, S. A. Sawant, S. Mishra, M. Nakamura, A. Mitra, S. Mashetty, and C. Baral · 2023
Later among the works it cites.
Continual instruction tuning for large multimodal models
J. He, H. Guo, M. Tang, and J. Wang · 2023
Later among the works it cites.
Parameter-level soft-masking for continual learning
T. Konishi, M. Kurokawa, C. Ono, Z. Ke, G. Kim, and B. Liu · 2023
Later among the works it cites.
Lora fine-tuning efficiently undoes safety training in llama 2-chat 70b
S. Lermen, C. Rogers-Smith, and J. Ladish · 2023
Later among the works it cites.
SPHINX: the joint mixing of weights, tasks, and visual embeddings for multi-modal large language models
Z. Lin, C. Liu, R. Zhang, P. Gao, L. Qiu, H. Xiao, H. Qiu, C. Lin, W. Shao, K. Chen, J. Han, S. Huang, Y. Zhang, X. He, H. Li, and Y. Qiao · 2023
Later among the works it cites.
Moelora: An moe-based parameter efficient fine-tuning method for multi-task medical applications
Q. Liu, X. Wu, X. Zhao, Y. Zhu, D. Xu, F. Tian, and Y. Zheng · 2023
Later among the works it cites.
G-eval: NLG evaluation using gpt-4 with better human alignment
Y. Liu, D. Iter, Y. Xu, S. Wang, R. Xu, and C. Zhu · 2023
Later among the works it cites.
Unified-io 2: Scaling autoregressive multimodal models with vision, language, audio, and action
J. Lu, C. Clark, S. Lee, Z. Zhang, S. Khosla, R. Marten, D. Hoiem, and A. Kembhavi · 2023
Later among the works it cites.
An empirical study of catastrophic forgetting in large language models during continual fine-tuning
Y. Luo, Z. Yang, F. Meng, Y. Li, J. Zhou, and Y. Zhang · 2023
Later among the works it cites.
GPT-4 technical report
OpenAI · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model, 2023
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample · 2023
Later among the works it cites.
Trace: A comprehensive benchmark for continual learning in large language models
X. Wang, Y. Zhang, T. Chen, S. Gao, S. Jin, X. Yang, Z. Xi, R. Zheng, Y. Zou, T. Gui, et al · 2023
Later among the works it cites.
Investigating the catastrophic forgetting in multimodal large language models
Y. Zhai, S. Tong, X. Li, M. Cai, Q. Qu, Y. J. Lee, and Y. Ma · 2023
Later among the works it cites.
Citb: A benchmark for continual instruction tuning
Z. Zhang, M. Fang, L. Chen, and M.-R. Namazi-Rad · 2023
Later among the works it cites.
A survey of large language models
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, Y. Du, C. Yang, Y. Chen, Z. Chen, J. Jiang, R. Ren, Y. Li, X. Tang, Z. Liu, P. Liu, J. Nie, and J. Wen · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny · 2023
Later among the works it cites.
Visual instruction tuning
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2024
Closest in time.