Fetching the paper…
Reading the bibliography…
As Multimodal Large Language Models (MLLMs) grow in size, adapting them to specialized tasks becomes increasingly challenging due to high computational and memory demands.
“Microsoft coco: Common objects in context,”
Tsung-Yi Lin, Michael Maire, et al., · 2014
Earlier work this paper cites.
“Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,”
Bryan A Plummer, Wang, et al., · 2015
Earlier work this paper cites.
“Learning multiple visual domains with residual adapters,”
Sylvestre-Alvise Rebuffi, Hakan Bilen, et al., · 2017
Earlier work this paper cites.
“Multimodal machine learning: A survey and taxonomy,”
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency, · 2018
Earlier work this paper cites.
“Parameter-efficient transfer learning for NLP,”
Neil Houlsby, Andrei Giurgiu, et al., · 2019
Earlier work this paper cites.
“VL-BERT: pre-training of generic visual-linguistic representations,”
Weijie Su, Xizhou Zhu, Yue Cao, et al., · 2020
Earlier work this paper cites.
“Automated crisis content categorization for covid-19 tweet streams,”
Zijun Long and Richard Mccreadie, · 2021
Earlier work this paper cites.
“Align before fuse: Vision and language representation learning with momentum distillation,”
Junnan Li, Ramprasaath Selvaraju, et al., · 2021
Earlier work this paper cites.
“Scaling up visual and vision-language representation learning with noisy text supervision,”
Chao Jia, Yinfei Yang, et al., · 2021
Earlier work this paper cites.
“Learning transferable visual models from natural language supervision,”
Alec Radford, Jong Wook Kim, et al., · 2021
Earlier work this paper cites.
“Lora: Low-rank adaptation of large language models,”
Edward J. Hu, Yelong Shen, et al., · 2022
Cited alongside, same era.
“VL-ADAPTER: parameter-efficient transfer learning for vision-and-language tasks,”
Yi-Lin Sung, Jaemin Cho, and Mohit Bansal, · 2022
Cited alongside, same era.
“Adaptformer: Adapting vision transformers for scalable visual recognition,”
Shoufa Chen, Chongjian Ge, Zhan Tong, et al., · 2022
Cited alongside, same era.
“Is multi-modal data key for crisis content categorization on social media?,”
Zijun Long and Richard Mccreadie, · 2022
Cited alongside, same era.
“Efficient adapter transfer of self-supervised speech models for automatic speech recognitio,”
Bethan Thomas, Samuel Kessler, et al., · 2022
Cited alongside, same era.
“An adapter based pre-training for efficient and scalable self-supervised speech representation learning,”
“Crisisvit: A robust vision transformer for crisis image classification,”
Zijun Long, Richard Mccreadie, and Imran Muhammad, · 2023
Closest in time.
“Robollm: Robotic vision tasks grounded on multimodal large language models,”
Zijun Long, George Killick, et al., · 2023
Closest in time.
“Large multi-modal encoders for recommendation,”
Zixuan Yi, Zijun Long, Iadh Ounis, et al., · 2023
Closest in time.
“Elucidating and overcoming the challenges of label noise in supervised contrastive learning,”
Zijun Long, George Killick, Lipeng Zhuang, et al., · 2023
Closest in time.
“Vppt: Visual pre-trained prompt tuning framework for few-shot image classification,”
Zhao Song, Ke Yang, Naiyang Guan, et al., · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Samuel Kessler, Bethan Thomas, et al., · 2022
Cited alongside, same era.
“Vlmo: Unified vision-language pre-training with mixture-of-modality-experts,”
Hangbo Bao, Wenhui Wang, et al., · 2022
Cited alongside, same era.
“BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,”
Junnan Li, Dongxu Li, Silvio Savarese, et al., · 2023
Cited alongside, same era.
“Image as a foreign language: Beit pretraining for all vision and vision-language tasks,”
Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, and et al., · 2023
Cited alongside, same era.
“Lit: Zero-shot transfer with locked-image text tuning,”
Xiaohua Zhai, Xiao Wang, Basil Mustafa, et al.,
Cited in the paper.
“Parameter-efficient transfer learning of pre-trained transformer models for speaker verification using adapters,”
Junyi Peng, Themos Stafylakis, Rongzhi Gu, et al., · 2023
Closest in time.
“Using adapters to overcome catastrophic forgetting in end-to-end automatic speech recognition,”
Steven Vander Eeckt and Hugo Van hamme, · 2023
Closest in time.
“Adapted multimodal bert with layer-wise fusion for sentiment analysis,”
Odysseas S. Chlapanis, Georgios Paraskevopoulos, et al., · 2023
Closest in time.
“Lacvit: A label-aware contrastive training framework for vision transformers,”
Zijun Long, Zaiqiao Meng, et al., · 2024
Closest in time.